Prompt
How do I set up alerts for LLM timeouts and error spikes?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
To set up alerts for LLM timeouts and error spikes, you want to monitor a few key signals, then define thresholds or anomaly rules around them.
1) Track the right metrics
At minimum, emit metrics for:
- Request count
- Error count
- Timeout count
- Latency (p50/p95/p99)
- Success rate
- Retry count
- Provider-specific errors (rate limits, 5xx, auth failures)
If you’re using an LLM SDK or gateway, make sure it logs:
- model name
- provider
- endpoint
- request status
- error type
- duration
- timeout flag
2) Define useful alert conditions
Common alerts:
Timeout alert
Trigger when:
- timeout rate > X% over Y minutes, or
- timeout count > N in Y minutes
Example:
timeouts / total_requests > 0.05for 5 minutes- or
timeouts > 20in 10 minutes
Error spike alert
Trigger when:
- error rate > X% over Y minutes
- 5xx or upstream errors exceed threshold
- a specific error type spikes
Example:
errors / total_requests > 0.02for 10 minutesprovider_5xx > 10in 5 minutes
Latency alert
Often timeouts are preceded by latency creep:
- p95 latency > threshold
- p99 latency rising rapidly
Example:
p95_latency > 8sfor 15 minutes
3) Use a moving window
Avoid alerting on single failures. Use rolling windows:
- 5 minutes for fast incidents
- 15 minutes for broader stability issues
This reduces noise and catches sustained problems.
4) Separate alerts by severity
A good structure:
- Warning: timeout/error rate elevated
- Critical: severe outage or sustained spike
Example:
- Warning: error rate > 1% for 10m
- Critical: error rate > 5% for 5m
5) Alert on per-model or per-provider issues
LLM incidents are often scoped:
- one model
- one region
- one provider
- one tenant/customer
So create labels/dimensions and alert on:
model=gpt-4.1provider=openairegion=us-east-1
6) Use logs for diagnosis, metrics for alerting
- Metrics: trigger alerts
- Logs/traces: debug root cause
Make sure each failed request logs:
- request ID
- prompt size / token count
- error message
- timeout duration
- retry attempts
7) Example alert logic
Pseudo-code:
timeout_rate = timeouts_last_5m / requests_last_5m
error_rate = errors_last_5m / requests_last_5m
if timeout_rate > 0.05:
page("LLM timeout spike")
if error_rate > 0.02:
page("LLM error spike")
8) Common tools
Depending on your stack:
- Datadog: metric monitors + logs
- Prometheus/Grafana: alert rules
- CloudWatch: metric alarms
- New Relic: NRQL alerts
- OpenTelemetry: instrumentation, then export to your monitoring backend
9) Best practices
- Alert on rates, not just raw counts
- Add minimum request volume to avoid noisy alerts
- Exclude expected failures like user cancellations
- Route alerts to Slack/PagerDuty/email
- Include a dashboard link in the alert
10) Example Prometheus alert rules
- alert: LLMTimeoutSpike
expr: sum(rate(llm_timeouts_total[5m])) / sum(rate(llm_requests_total[5m])) > 0.05
for: 5m
labels:
severity: critical
annotations:
summary: "LLM timeout spike"
- alert: LLMErrorSpike
expr: sum(rate(llm_errors_total[5m])) / sum(rate(llm_requests_total[5m])) > 0.02
for: 10m
labels:
severity: warning
annotations:
summary: "LLM error spike"
If you want, I can also give you:
- a Datadog setup,
- a Prometheus/Grafana example, or
- a Python/OpenTelemetry instrumentation snippet.