Prompt
I need LLM monitoring with alerting for latency, timeout, and token spend spikes
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Here’s a practical way to set up LLM monitoring with alerting for latency, timeouts, and token spend spikes.
What to monitor
1) Latency
Track:
- p50 / p95 / p99 latency
- latency by:
- model
- endpoint / feature
- tenant / customer
- request type (chat, embedding, tool call, etc.)
Alert on:
- sustained p95 latency above threshold
- sudden regression compared with baseline
- per-model or per-tenant latency anomalies
2) Timeouts
Track:
- request timeout count
- timeout rate = timeouts / total requests
- upstream provider timeouts vs app timeouts
- retries and retry exhaustion
Alert on:
- timeout rate above threshold
- burst of timeouts in a short window
- increase in retries leading to degraded UX
3) Token spend
Track:
- input tokens
- output tokens
- total tokens
- cost estimates by model
- tokens by tenant, feature, user, and prompt version
Alert on:
- sudden token spike vs historical baseline
- daily spend approaching budget
- unusual token-per-request increase
- high-cost model usage surge
Recommended metrics
A solid minimum set:
llm_requests_totalllm_request_latency_msllm_request_timeouts_totalllm_request_errors_totalllm_input_tokens_totalllm_output_tokens_totalllm_total_tokens_totalllm_estimated_cost_usd_total
Useful labels:
modelproviderendpointtenantenvironmentstatusreason
Avoid high-cardinality labels like raw user IDs unless you aggregate carefully.
Alert examples
Latency alert
Trigger if:
- p95 latency > 2s for 5 minutes
- or latency increases > 50% over 1 hour baseline
Example:
- Warning: p95 > 1500 ms for 10 min
- Critical: p95 > 3000 ms for 5 min
Timeout alert
Trigger if:
- timeout rate > 2% for 5 minutes
- or > 10 timeouts in 1 minute
Token spend spike alert
Trigger if:
- tokens per minute > 2x trailing 24h average
- or daily estimated cost crosses 80% of budget
- or one tenant exceeds 3x normal usage
Best practice architecture
Instrumentation
Add telemetry in your LLM wrapper or middleware:
- start timer at request start
- count tokens from prompt + completion
- record status, retries, timeout reason
- attach model/provider metadata
Storage / dashboards
Send metrics to one of:
- Prometheus + Grafana
- Datadog
- OpenTelemetry + backend
- CloudWatch / Azure Monitor / GCP Monitoring
Alerts
Route alerts to:
- Slack / Teams
- PagerDuty / Opsgenie
- email for budget warnings
Dashboards
Create panels for:
- request volume
- latency percentiles
- timeout rate
- token usage rate
- cost by model/tenant
- top outlier requests
A good starting threshold set
If you don’t have baselines yet:
-
Latency
- warn: p95 > 1.5s for 10 min
- critical: p95 > 3s for 5 min
-
Timeouts
- warn: timeout rate > 1%
- critical: timeout rate > 3%
-
Token spend
- warn: 24h spend > 75% of budget
- critical: 24h spend > 90% of budget
- spike alert: > 2x 7-day rolling average for 15 min
Example alerting logic
IF p95_latency_5m > threshold
ALERT "LLM latency high"
IF timeout_rate_5m > threshold
ALERT "LLM timeouts increasing"
IF token_usage_15m > 2 * avg_token_usage_24h
ALERT "LLM token spend spike"
If you want a production-ready setup
I can help you with any of these next:
- Prometheus/Grafana metrics and alert rules
- Datadog monitors
- OpenTelemetry instrumentation in Python/Node/Go
- A complete LLM observability schema
- Budget guardrails and anomaly detection logic
If you tell me your stack, I can generate the exact config and code.