Prompt
How do I set up alerts for latency, token spend, and failure spikes in an LLM app?
Latest observation
Here’s a practical way to set up alerts for an LLM app so you catch problems early without getting paged for noise.
1) Track the right metrics
At minimum, emit these per request:
Latency
request_latency_ms- Break out by:
- model
- endpoint / route
- customer / tenant
- success vs error
- streaming vs non-streaming
Useful aggregations:
- p50, p95, p99 latency
- average latency
- tail latency by model/version
Token spend
prompt_tokenscompletion_tokenstotal_tokensestimated_cost_usd
Also track:
- tokens per request
- tokens per user / org / day
- spend by model, route, and tenant
Failures
request_counterror_counttimeout_countrate_limit_countinvalid_response_counttool_call_failure_count
Calculate:
- error rate =
errors / total_requests - timeout rate
- 429 rate
- retry rate
2) Use alert thresholds that match the metric
A good alerting setup usually has:
- absolute thresholds for hard limits
- relative anomaly alerts for spikes
- burn-rate alerts for sustained issues
Latency alerts
Examples:
- p95 latency > 2s for 10 minutes
- p99 latency > 5s for 5 minutes
- latency increased 2x vs 7-day baseline
A good pattern:
- Warning: p95 above threshold for 10–15 min
- Critical: p99 above threshold or big regression vs baseline
Token spend alerts
Examples:
- Daily spend > 80% of budget
- Hourly spend > 2x normal rate
- Tokens per request increased 50% vs baseline
- Completion tokens spike on a specific route/model
If you have a budget:
- Alert at 50%, 80%, 90%, and 100%
- Add anomaly alerts for sudden growth
Failure spike alerts
Examples:
- Error rate > 5% for 5 minutes
- Timeout rate > 2% for 10 minutes
- 429 rate > 1% for 5 minutes
- Success rate drops below 95%
For production, a classic SRE-style alert is:
5xx/total > threshold over rolling window- burn-rate alert for your error budget
3) Alert on ratios, not just counts
Counts alone are noisy. Prefer:
error_rate = errors / requeststimeout_rate = timeouts / requeststoken_cost_per_requestlatency_p95compared to baseline
This avoids false alarms when traffic is low or high.
Example:
- Don’t alert on “100 failures”
- Do alert on “failure rate above 3% for 10 minutes, with at least 200 requests”
That minimum volume guard is important.
4) Add baseline/anomaly detection for token spend and latency
LLM apps often have natural variability, so static thresholds aren’t enough.
Use:
- moving average
- day-over-day comparison
- week-over-week comparison
- seasonal baselines by hour/day
Good anomaly examples:
- token spend today is 1.8x the same hour last week
- p95 latency is 2 standard deviations above normal
- completion length jumped after a prompt/model change
5) Segment alerts by model, route, and tenant
Most LLM incidents are localized.
Break alerts down by:
- model version
- prompt template
- feature flag
- customer / tenant
- geography / region
- provider (OpenAI, Anthropic, local model, etc.)
This helps you avoid broad alerts when only one route is broken.
6) Create a small alert matrix
A simple starting point:
Latency
- Warning: p95 > 1500 ms for 10 min
- Critical: p99 > 4000 ms for 5 min
Spend
- Warning: daily spend > 75% of budget
- Critical: daily spend > 90% of budget
- Spike: spend/hour > 2x 7-day hourly baseline
Failures
- Warning: error rate > 2% for 10 min
- Critical: error rate > 5% for 5 min
- Spike: timeout rate > 1% for 5 min
- Spike: 429 rate > 1% for 5 min
Adjust the numbers based on your traffic and business tolerance.
7) Instrument before you alert
Alerts are only useful if the underlying data is reliable.
Log at least:
- request ID
- user/tenant ID
- model name/version
- prompt template version
- token counts
- latency breakdown:
- queue time
- provider latency
- tool latency
- post-processing latency
- response status / error type
If possible, also capture:
- retries
- cached vs uncached requests
- streaming duration
- prompt length and output length
8) Use dashboards plus alerts
Before creating alerts, build dashboards for:
- total requests
- latency percentiles
- token spend over time
- errors by type
- spend by model/tenant
- top slow routes
- top expensive prompts
Then create alerts from the same metrics.
9) Route alerts by severity
Example:
- Warning → Slack / email
- Critical → PagerDuty / phone
- Business spend breach → finance/ops channel too
Also include:
- current value
- baseline
- impacted route/model
- link to dashboard
- recent deploy or prompt change if available
10) Reduce noisy alerts
Best practices:
- require N minutes over threshold
- require minimum traffic
- deduplicate similar alerts
- suppress during planned maintenance
- use different thresholds for dev/staging/prod
- alert only on user-facing paths, not every background job
Example alert rules
Latency
p95(request_latency_ms) > 2000 for 10mp99(request_latency_ms) > 5000 for 5m
Token spend
sum(estimated_cost_usd) over 1h > 2x avg of same hour last 7 daysdaily_spend > budget * 0.8
Failure
error_rate > 0.05 for 5m and request_count > 200timeout_rate > 0.02 for 10m
Suggested stack
Common implementation:
- Metrics: Prometheus, Datadog, OpenTelemetry
- Dashboards: Grafana, Datadog dashboards
- Alerts: PagerDuty, Opsgenie, Slack
- Logs/traces: OpenTelemetry + your logging backend
If you want, I can also give you:
- a Prometheus alert rule set for these metrics,
- a Datadog monitor setup, or
- a Python/Node instrumentation example for an LLM app.