Prompt
I'm building an API product and need response code and latency alerts
Latest observation
Oct 3, 2026
OpenAI APIWeb search: off
Here’s a practical way to set up response code and latency alerts for an API product.
1) What to alert on
Response code alerts
Track by status family and endpoint:
- 5xx rate: server errors
- 4xx rate: client errors, but usually only alert on unusual spikes
- Specific codes: 401/403, 429, 404, 502/503/504 depending on your API
- Per-endpoint error rate: important so one bad route doesn’t hide in aggregate
Good baseline alerts:
- 5xx rate > 1% for 5 minutes
- 5xx count > N per minute
- 429 rate spikes if rate limiting or abuse matters
- Critical endpoint returns > X errors/min
Latency alerts
Use percentiles, not averages:
- p50 for typical traffic
- p95 for user experience
- p99 for tail latency
Good baseline alerts:
- p95 latency > target for 5 minutes
- p99 latency > 2x normal baseline
- Sustained increase vs. trailing 1h/24h baseline
Example targets:
- Public API: p95 < 300–500 ms
- Internal API: p95 < 100–200 ms
- Very fast services: p95 < 50 ms
2) Best alerting signals
Use these metrics:
request_counterror_count{status_code, endpoint}latency_ms_bucketor histogramlatency_p95,latency_p99availabilityor success rate
Also slice by:
- endpoint
- method
- region
- customer tier
- auth type
3) Alert design principles
To avoid noisy alerts:
- Alert on rates, not raw counts alone
- Use rolling windows like 5m, 10m
- Add minimum traffic thresholds so low-volume endpoints don’t trigger constantly
- Prefer multi-condition alerts:
- “5xx > 2% AND requests > 100 in 5m”
- Separate warning and critical thresholds
Example:
- Warning: p95 > 400ms for 10m
- Critical: p95 > 800ms for 5m
4) Example alert rules
Response code
- Critical:
5xx_rate > 2% over 5m503_count > 10 in 5m
- Warning:
5xx_rate > 1% over 10m4xx_rate > 10% over 10mif unusual
Latency
- Critical:
p95_latency > 1s over 5mp99_latency > 2s over 5m
- Warning:
p95_latency > 500ms over 10m
5) If you use Prometheus/Grafana
Example PromQL-style patterns:
5xx rate
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
p95 latency
histogram_quantile(
0.95,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)
Endpoint-specific error rate
sum(rate(http_requests_total{status=~"5..", route="/v1/payments"}[5m]))
/
sum(rate(http_requests_total{route="/v1/payments"}[5m]))
6) Recommended operational setup
- Dashboard: traffic, error rate, p95/p99 latency, top failing endpoints
- Alerts: route to Slack/PagerDuty/email
- Runbooks: for each critical alert, define likely causes and first checks
- SLOs: tie alerts to user experience, not just raw metrics
7) Simple starting policy
If you want a fast starting point:
- Alert when 5xx > 1% over 5m
- Alert when p95 latency > 500ms over 5m
- Alert when p99 latency > 1.5s over 5m
- Add per-endpoint alerts for critical routes
- Add a minimum request volume threshold
If you want, I can give you:
- a Prometheus/Grafana alert rule pack,
- an AWS CloudWatch setup, or
- a vendor-neutral alerting spec you can hand to your team.