Prompt
How do I find which service added latency in a trace
Latest observation
Sep 30, 2026
OpenAI APIWeb search: off
To find which service added latency in a trace, look for where the time increases between spans and which span has the largest duration or largest gap before/after it.
What to check
-
Total trace duration
- Compare the root span’s duration to child spans.
- A slow root with normal children usually means time spent in waiting, network, queueing, or uninstrumented code.
-
Span durations
- Find the span with the longest duration.
- That service is often contributing the most latency.
-
Exclusive time vs inclusive time
- A span may look long because it contains child calls.
- The real culprit is often the service with high exclusive/self time.
-
Gaps between spans
- If there’s a delay between parent and child spans, it may indicate:
- queue delays
- network latency
- thread pool saturation
- async scheduling delays
- missing instrumentation
- If there’s a delay between parent and child spans, it may indicate:
-
Critical path
- In distributed traces, the slowest service may not be the one with the most spans.
- Follow the path that determines the overall request time.
Practical method
- Open the trace waterfall/timeline view.
- Sort mentally by:
- longest span
- largest idle gap
- spans on the critical path
- Identify the service where latency first starts increasing and see whether downstream services are actually slow or just waiting.
Example
If a request takes 2s total:
- API Gateway: 50ms
- Auth: 40ms
- Orders service: 1.6s
- DB query inside Orders: 1.5s
Then the DB query is the likely latency source, and the Orders service is where it shows up.
If you’re using common tracing tools
- Jaeger / Zipkin / Honeycomb / Datadog / New Relic:
- inspect the span tree
- check span duration, child spans, and gaps
- use “critical path” or “service map” views if available
If you want, I can also show you how to identify the slow service from a specific trace screenshot or span JSON.