Prompt
How do I correlate logs, metrics, and traces for troubleshooting?
Latest observation
Oct 5, 2026
OpenAI APIWeb search: off
A good way to troubleshoot is to treat logs, metrics, and traces as three views of the same incident:
- Metrics tell you that something is wrong and roughly when / how much
- Traces tell you where in the request path the problem is happening
- Logs tell you why it happened, with details and context
A practical workflow
1) Start with metrics
Use metrics to identify the symptom:
- Increased latency
- Higher error rate
- Dropped throughput
- Saturation (CPU, memory, disk, queue depth)
This helps narrow the time window and service/component.
Example:
5xx rate spiked at 14:32p95 latency doubled on checkout-serviceDB pool exhaustion started
2) Jump from metrics to traces
Once you know the affected service and timeframe, inspect traces for:
- Slow spans
- Error spans
- Unexpected downstream calls
- Retry storms
- Hot paths or bottlenecks
Look for:
- A request ID / trace ID tied to the bad request
- Which child span is dominating latency
- Whether the issue is in your service or a dependency
3) Pivot from traces to logs
Use the trace ID, span ID, or request ID to pull related logs:
- Application logs from the same request
- Dependency logs around the same timestamp
- Infrastructure logs if needed
This often reveals:
- Exceptions
- Validation failures
- Timeouts
- Misconfigurations
- Rate limiting or authentication issues
The key: shared identifiers
Correlation works best when all three signals share common fields:
timestampservice.namehost/pod/containerenvironmenttrace_idspan_idrequest_id/correlation_iduser_id/tenant_idif relevant
If possible, make sure:
- Logs include
trace_id - Traces include service and operation names
- Metrics are labeled by service, route, status code, and environment
A simple example
Suppose:
- Metrics show checkout latency rising
- Traces show most time spent waiting on
payment-api - Logs for those trace IDs show
timeout connecting to payment gateway
That gives you the full story:
- Metric: latency spike
- Trace: slow downstream dependency
- Logs: exact timeout/error message
Best practices
- Use structured logging (JSON, key-value logs)
- Propagate distributed trace context via headers
- Keep metric labels low-cardinality to avoid explosion
- Set alerts on SLIs/SLOs, then investigate with traces/logs
- Standardize timestamps and time zones (prefer UTC)
- Centralize observability data in one tool or linked tools
A good troubleshooting pattern
- Alert fires from metrics
- Filter traces for the impacted service and time
- Find slow/error spans
- Copy trace ID into log search
- Confirm root cause from log details
- Validate with additional metrics
If you want, I can also give you:
- a step-by-step example in Kubernetes/OpenTelemetry
- a tool-specific workflow for Datadog, Grafana, Splunk, New Relic, or Elastic
- a template for logging and trace correlation fields