Prompt
Need to track API latency and bottlenecks
Latest observation
Oct 5, 2026
OpenAI APIWeb search: off
To track API latency and find bottlenecks, focus on measurements at each layer of the request path, not just end-to-end timing.
What to track
1. Core latency metrics
Measure these for every endpoint:
- Request duration: total time from arrival to response
- P50 / P95 / P99 latency: median and tail latency
- Error rate: 4xx/5xx, timeouts
- Throughput: requests/sec
- Payload size: request/response sizes
2. Break latency into stages
Instrument each request with spans or timers:
- Network ingress
- Auth / middleware
- Request parsing / validation
- Business logic
- DB queries
- Cache calls
- External API calls
- Response serialization / egress
This helps identify whether the bottleneck is in:
- application code
- database
- downstream services
- serialization
- network
3. Dependency metrics
For each dependency track:
- latency per call
- error rate
- retry count
- timeout count
- connection pool saturation
- queue time / wait time
4. Infrastructure metrics
Also monitor:
- CPU
- memory
- GC pauses
- thread/worker saturation
- event loop lag
- open connections
- disk I/O
- network I/O
Best tools / patterns
Distributed tracing
Use tracing to see a request across services.
- OpenTelemetry is the common standard
- Export to Jaeger, Tempo, Datadog, New Relic, Honeycomb, etc.
Metrics
Use time-series metrics for dashboards and alerts.
- Prometheus + Grafana is a common stack
- Track latency histograms, not just averages
Logs
Use structured logs with:
- request ID / trace ID
- endpoint
- duration
- dependency timings
- error details
How to pinpoint bottlenecks
1. Compare endpoint latency to dependency spans
If endpoint latency is high and one DB span is also high, the DB is likely the issue.
2. Look at tail latency
P95/P99 spikes often reveal:
- lock contention
- slow queries
- cold caches
- pool exhaustion
- noisy neighbors
- retries
3. Correlate with saturation
If latency rises with CPU, memory, or pool usage, you likely have a capacity bottleneck.
4. Inspect slow queries and hot paths
Common causes:
- missing DB indexes
- N+1 queries
- excessive serialization
- chatty internal calls
- large payloads
- synchronous work on critical path
Practical setup
A good minimal setup:
- Add request timing middleware
- Add tracing with OpenTelemetry
- Export metrics to Prometheus
- Create dashboards for p50/p95/p99, errors, throughput
- Alert on latency and saturation
- Sample traces for slow requests
Example metrics to alert on
- P95 latency > threshold for 5–10 min
- Error rate > 1%
- DB query latency spike
- pool utilization > 80–90%
- timeout rate increasing
If you want, I can also give you:
- a language-specific implementation (Python, Node, Go, Java, etc.), or
- a full observability architecture for APIs.