Prompt
I'm building an LLM product and need alerts for latency, cost, and failure spikes
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Here’s a practical alerting setup for an LLM product focused on latency, cost, and failures.
1) What to monitor
Latency
Track:
- p50 / p95 / p99 end-to-end latency
- time to first token (TTFT) if streaming
- model inference latency
- queue/wait time before inference
- tool/API latency if your app uses tools
Cost
Track:
- tokens per request
- cost per request
- cost per user / tenant / feature
- daily spend
- spend by model
- input vs output token mix
Failures
Track:
- request error rate
- timeout rate
- rate-limit / quota errors
- provider 5xx / 4xx errors
- retry rate
- fallback rate
- guardrail / moderation blocks
- empty / malformed outputs
- schema validation failures
2) Good alert thresholds to start with
Use both absolute thresholds and change-from-baseline thresholds.
Latency alerts
- Critical: p95 latency > 2x baseline for 10–15 min
- Warning: p95 latency > 1.5x baseline for 15 min
- Critical: TTFT > 3s for >5% of requests
- Warning: p99 latency exceeds your SLO for 10 min
Cost alerts
- Critical: daily spend projected to exceed budget by 20%
- Warning: spend rate up 30–50% vs baseline over 1h
- Critical: tokens/request up 2x baseline
- Warning: output tokens spike by 30%+ over baseline
Failure alerts
- Critical: error rate > 5% for 5 min
- Warning: error rate > 1% for 15 min
- Critical: timeout rate > 2% for 10 min
- Critical: provider 429/5xx spike above baseline by 3x
- Warning: fallback usage doubles vs baseline
3) Alert design best practices
Prefer symptom + cause alerts
Example:
- Symptom: “p95 latency is high”
- Cause: “provider timeouts increased” This helps avoid noisy alerts.
Use burn-rate alerts for SLOs
If you have SLOs like:
- 99% of requests under 3s
- 99.5% success rate
Alert on error budget burn:
- Fast burn: 14x burn over 5 min
- Slow burn: 2x burn over 1h
This is more reliable than static thresholds.
Segment by dimension
Break alerts down by:
- model
- tenant/customer
- region
- endpoint/feature
- prompt template/version
This helps catch one bad release or one noisy customer.
4) Recommended dashboard metrics
A simple dashboard should show:
- Request volume
- p50/p95/p99 latency
- TTFT
- Error rate
- Timeout rate
- Retry rate
- Token usage/request
- Spend/day and spend/hour
- Cost/request
- Provider status / rate limits
- Fallback rate
5) Example alerts to implement
Latency
- If p95 latency > 1.5x 7-day baseline for 15 min → warning
- If p95 latency > 2x baseline or > SLO for 10 min → critical
Cost
- If projected daily spend > budget by 20% → critical
- If avg cost/request > 1.3x baseline for 1h → warning
- If output tokens/request > 2x baseline for 15 min → warning
Failures
- If overall error rate > 1% for 15 min → warning
- If overall error rate > 5% for 5 min → critical
- If timeout rate > 2% for 10 min → critical
- If 429s or 5xx from provider > 3x baseline → critical
6) Operational suggestions
- Page only on user-impacting issues.
- Send non-urgent spikes to Slack/email.
- Include links in alerts to:
- traces
- logs
- model/version
- tenant breakdown
- Add auto-triage metadata:
- prompt version
- model version
- provider
- region
- rollout status
7) If you want a simple initial rule set
Start with:
- p95 latency > 2x baseline for 10 min
- error rate > 5% for 5 min
- daily spend forecast > budget by 20%
- timeout rate > 2% for 10 min
- 429/5xx spikes > 3x baseline
If you want, I can turn this into:
- a Datadog/Prometheus alert config
- a Grafana dashboard spec
- or a Python/SQL monitoring schema for your LLM app.
Brands