Prompt
How do I track API latency and error rates in real time?
Latest observation
To track API latency and error rates in real time, you want three things: good instrumentation, a metrics pipeline, and alerts/dashboards.
1) Instrument the API
Capture request/response data for every request:
- Latency: time from request start to response end
- Status code: 2xx / 4xx / 5xx
- Error type: timeout, validation error, dependency failure, exception
- Endpoint: route/path
- Method: GET/POST/etc.
- Correlation/request ID: to trace individual requests
At minimum, emit a metric per request like:
api_request_duration_msapi_request_totalapi_request_errors_total
2) Use metrics, not logs, for real-time tracking
Metrics are best for live monitoring because they can be aggregated quickly.
Common latency metrics:
- Average latency (useful, but can hide spikes)
- Percentiles: p50, p90, p95, p99
- Apdex if you want a user-experience score
Common error metrics:
- Error rate =
5xx requests / total requests - Also track 4xx separately so client errors don’t inflate server error rate
Example formulas:
- Latency p95: 95% of requests complete faster than this value
- 5xx error rate:
count(status >= 500) / count(all requests)
3) Export metrics to a monitoring system
Popular setups:
- Prometheus + Grafana
- Datadog
- New Relic
- CloudWatch / Azure Monitor / GCP Cloud Monitoring
If you use Prometheus:
- Expose a
/metricsendpoint - Scrape it every few seconds
- Visualize in Grafana
- Alert with Alertmanager
4) Add distributed tracing
For real-time diagnosis, add tracing with:
- OpenTelemetry
- Jaeger / Tempo / Zipkin / Datadog APM
Tracing helps answer:
- Which downstream service caused the latency?
- Is the slowdown in your app, database, or third-party API?
5) Build useful dashboards
Show:
- Request volume per minute
- p50/p95/p99 latency per endpoint
- 4xx and 5xx error rates
- Top slow endpoints
- Dependency latency
- Saturation metrics: CPU, memory, DB connections, queue depth
6) Set alerts
Alert on symptoms, not just raw numbers:
- Error rate > 1% for 5 minutes
- p95 latency > 500 ms for 10 minutes
- Sudden drop in throughput
- Spike in 5xx errors
Use multi-window alerts to reduce noise.
7) Sample implementation idea
If you’re using Python with Prometheus, you might record:
- a histogram for latency
- a counter for total requests
- a counter for errors
Then your dashboard can compute:
- request rate
- error rate
- p95 latency
8) Best practices
- Track metrics per endpoint and per status class
- Avoid high-cardinality labels like user IDs or full URLs with query strings
- Use histograms for latency, not just averages
- Separate application errors from dependency errors
- Keep dashboards and alerts focused on user impact
If you tell me your stack (e.g., Node.js, Python, Go, Java, AWS, Kubernetes), I can give you a concrete setup and code example.