Prompt
I need a production monitoring setup for token usage, timeouts, and errors
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Here’s a solid production monitoring setup for token usage, timeouts, and errors for an LLM/API-based system.
1) What to monitor
Token usage
Track per request and aggregate:
- prompt_tokens
- completion_tokens
- total_tokens
- tokens/sec
- tokens per request by endpoint / model / tenant
- cost estimate if pricing varies by model
Timeouts / latency
Track:
- request latency p50 / p95 / p99
- model inference time
- queue/wait time
- timeout count
- retry count
- stream duration if using streaming responses
Errors
Track:
- HTTP status codes
- provider/API errors
- rate limit errors
- context length exceeded
- invalid request / schema errors
- internal exceptions
- retry exhaustion
- circuit breaker opens
2) Core metrics to implement
Counters
requests_totalrequests_failed_totaltimeouts_totalretries_totaltoken_prompt_totaltoken_completion_totaltoken_total_totalprovider_errors_total{type=...}rate_limit_errors_totalcontext_length_errors_total
Histograms
request_latency_secondsmodel_latency_secondsqueue_wait_secondsresponse_size_bytestokens_per_request
Gauges
in_flight_requestscurrent_queue_depthcircuit_breaker_stateworker_utilization
3) Logging structure
Use structured logs with these fields:
request_idtrace_iduser_idortenant_idmodelendpointstatuslatency_msprompt_tokenscompletion_tokenstotal_tokenserror_typeerror_messageretry_counttimeout_mscost_estimate
Example:
{
"level": "error",
"request_id": "req_123",
"trace_id": "trace_abc",
"tenant_id": "tenant_42",
"model": "gpt-4.1",
"endpoint": "/chat",
"status": 504,
"latency_ms": 30012,
"timeout_ms": 30000,
"total_tokens": 1824,
"error_type": "timeout",
"retry_count": 2
}
4) Tracing setup
Instrument one trace per request:
- API ingress
- auth/validation
- prompt building
- model call
- post-processing
- persistence / downstream calls
Add spans for:
- prompt assembly
- model request
- retries
- tool calls
- streaming
- fallback path
This helps separate:
- provider latency
- your app latency
- network latency
- queue delays
Use OpenTelemetry if possible.
5) Dashboards
Create a dashboard with:
Traffic
- requests/min
- active requests
- queue depth
Token usage
- total tokens over time
- tokens/request
- tokens by model
- tokens by tenant
- estimated spend
Reliability
- error rate
- timeout rate
- retry rate
- provider error breakdown
Performance
- p50/p95/p99 latency
- model latency
- queue wait time
Saturation
- worker utilization
- concurrency limits hit
- rate limit events
6) Alerts
Set alerts on:
- error rate > threshold
- timeouts spike
- p95 latency above SLO
- token usage spike vs baseline
- rate limit errors increase
- retry exhaustion
- queue depth growing
- cost anomaly
- provider outage indicators
Example alert thresholds:
- error rate > 2% for 5 min
- timeout rate > 1% for 5 min
- p95 latency > 10s for 10 min
- token usage > 3x daily baseline
- retry exhaustion > 20 in 10 min
7) Recommended stack
Simple, common stack
- OpenTelemetry for instrumentation
- Prometheus for metrics
- Grafana for dashboards/alerts
- Loki / ELK for logs
- Jaeger / Tempo for traces
If you use a managed platform
- Datadog
- New Relic
- Honeycomb
- Grafana Cloud
8) Best practices
- Tag metrics by model, tenant, endpoint, and environment
- Avoid high-cardinality labels like raw user IDs in Prometheus
- Sample logs/traces if traffic is high
- Record both client-side and server-side timeouts
- Distinguish soft failures (retries succeeded) from hard failures
- Track actual billed tokens, not just estimated tokens
- Add SLOs for:
- availability
- latency
- error rate
- cost per request
9) Minimal implementation plan
- Add structured logging
- Add OpenTelemetry traces
- Expose Prometheus metrics
- Create Grafana dashboard
- Add alert rules
- Build cost/token anomaly detection
- Add retry and timeout classification
10) If you want a ready-to-use schema
I can provide one of these next:
- Prometheus metrics definitions
- OpenTelemetry instrumentation example
- Grafana dashboard JSON
- Datadog monitor setup
- A full monitoring spec for LLM APIs
If you want, I can generate a production-ready setup for Python, Node.js, or Go.