Prompt
How do I troubleshoot API latency spikes and identify where the bottleneck is?
Latest observation
To troubleshoot API latency spikes, work from the outside in: first confirm when the spikes happen, then where time is spent (client, gateway, app, downstream services, DB, network), and finally correlate latency with a specific resource bottleneck.
1) Start with the shape of the problem
Answer these first:
- Is it all endpoints or only specific ones?
- Is it all traffic or only certain customers/regions/HTTP methods?
- Is it constant slow or spiky?
- Is it p50, p95, or p99 latency increasing?
- Does it happen at a specific time of day or during deploys?
Spikes limited to p95/p99 often point to:
- queueing
- GC pauses
- cold caches
- tail latency from downstream calls
- lock contention
- retry storms
2) Break latency into stages
Instrument or inspect the request path with timing for:
- DNS / TCP / TLS
- API gateway / load balancer
- application processing
- database queries
- cache calls
- external HTTP calls
- response serialization
- network egress
If you already have tracing, this is the fastest way to find the bottleneck. A distributed trace will show which span consumes time.
Good signals to look at
- High app time, low DB time → CPU, locks, serialization, code path issue
- High DB time → slow query, missing index, DB saturation, lock waits
- High external call time → dependency slowness, retries, timeout settings
- High queueing time → thread pool exhaustion, connection pool exhaustion, worker backlog
- High network time → packet loss, cross-zone/region traffic, DNS/TLS issues
3) Check metrics in parallel
Correlate latency spikes with:
- CPU utilization
- memory usage / GC
- thread pool saturation
- connection pool saturation
- request queue length
- disk I/O
- network I/O
- error rate / retries
- database QPS, locks, slow queries
- cache hit rate
- upstream/downstream latency
A common pattern is latency rising before errors do.
4) Use tracing to identify the slow span
If you have OpenTelemetry, Jaeger, Tempo, Datadog, Honeycomb, etc.:
- find a slow request sample
- inspect its critical path
- compare a slow trace to a normal one
- look for:
- long child spans
- repeated retries
- large gaps between spans
- time spent waiting on a pool or queue
If you don’t have tracing, add it. Without traces, you’re mostly guessing.
5) Inspect the most common bottlenecks
A. Database bottlenecks
Symptoms:
- app threads waiting on queries
- increased DB CPU/locks
- slow query logs showing same queries
- p99 spikes under load
What to check:
- query plans
- missing indexes
- N+1 queries
- connection pool exhaustion
- lock contention / deadlocks
- long transactions
- hot rows / hot partitions
B. Downstream service bottlenecks
Symptoms:
- one dependency dominates trace time
- retries amplify latency
- spikes align with that service’s incidents
What to check:
- dependency latency/error rate
- timeout values
- retry count and backoff
- circuit breaker behavior
- rate limiting / throttling
C. Application/runtime bottlenecks
Symptoms:
- high CPU or GC
- long pauses
- high context switching
- thread pool exhaustion
What to check:
- CPU profiling / flame graphs
- GC logs and pause time
- lock contention
- synchronous I/O on request threads
- inefficient serialization/deserialization
- memory pressure causing paging
D. Infrastructure/network bottlenecks
Symptoms:
- latency spikes across many endpoints at once
- region-specific issues
- elevated packet loss or retransmits
- load balancer or gateway saturation
What to check:
- LB/gateway metrics
- NIC saturation
- packet loss/retransmits
- cross-AZ/cross-region hops
- DNS lookup latency
- TLS handshake overhead
6) Compare good vs bad requests
Take one fast and one slow request and compare:
- same route?
- same customer/tenant?
- same payload size?
- same upstream/downstream calls?
- same code path?
- same shard/partition?
This often reveals:
- one tenant causing hot partitions
- large payloads
- cache misses
- a specific feature flag or code path
7) Check for hidden amplifiers
These often turn small issues into spikes:
- retries without jitter
- synchronized cron jobs
- cache stampedes
- thundering herd after deploy
- connection pool starvation
- autoscaling delays
- cold starts
- lock contention
- queue buildup
- noisy neighbors
8) Reproduce under controlled load
Use load testing to determine:
- max throughput before latency inflects
- whether latency rises gradually or suddenly
- which resource saturates first
Plot:
- latency vs RPS
- CPU vs RPS
- DB latency vs RPS
- queue length vs RPS
The first metric to saturate is usually your bottleneck.
9) A practical investigation workflow
- Confirm spike window and affected endpoints.
- Look at p50/p95/p99, error rate, throughput.
- Open a trace for a slow request.
- Identify the longest span.
- Check metrics for that component.
- Compare to a normal period.
- Inspect logs for timeouts, retries, pool exhaustion, GC, slow queries.
- Reproduce with load if needed.
- Fix the bottleneck.
- Add an alert and a dashboard to catch it early.
10) If you need a quick decision tree
- Slow span in DB? → query/index/locks/pool
- Slow span in another service? → dependency health/retries/timeouts
- No downstream span slow, but request slow? → app code, CPU, GC, locks
- Many requests slow at once? → infrastructure, saturation, queueing
- Only p99 spikes? → tail-latency causes: retries, contention, cold cache, noisy neighbor
11) Useful tools
- Distributed tracing: OpenTelemetry, Jaeger, Tempo, Datadog, Honeycomb
- APM: New Relic, AppDynamics, Datadog APM
- Profiling: pprof, async-profiler, Parca, Py-spy
- DB analysis: slow query logs, EXPLAIN, performance schema
- Metrics: Prometheus + Grafana
12) What to capture in an incident
At minimum:
- timestamp of spikes
- affected endpoints
- p50/p95/p99 latency
- error rate
- RPS
- CPU/memory/GC
- DB latency and slow queries
- downstream service latency
- queue lengths / pool usage
- trace samples
If you want, I can also give you:
- a step-by-step incident checklist,
- a Grafana dashboard template, or
- a tracing-based debugging flowchart.