Prompt
How do I trace a request through microservices in production using OpenTelemetry
Latest observation
To trace a request through microservices in production with OpenTelemetry, you want to ensure consistent context propagation, instrumentation in each service, and export to a tracing backend so you can follow a single request end-to-end.
1) Add OpenTelemetry instrumentation to every service
Each microservice should:
- create a span for incoming requests
- propagate the trace context to outgoing calls
- export traces to a backend
Typical instrumentation targets:
- HTTP servers/clients
- gRPC
- message queues / consumers / producers
- database calls
- background jobs
If you use a supported language/framework, prefer the OpenTelemetry auto-instrumentation or SDK libraries.
2) Propagate trace context between services
OpenTelemetry uses W3C Trace Context by default:
traceparenttracestate
When Service A calls Service B, the trace context must be injected into the outbound request headers. Service B extracts it and continues the same trace.
Example flow:
- Client → Service A: new trace starts
- Service A → Service B: context injected into headers
- Service B → Service C: same propagation
- All spans share the same
trace_id
3) Use a centralized exporter
Send spans from all services to an OpenTelemetry Collector or directly to a tracing system like:
- Jaeger
- Grafana Tempo
- Zipkin
- Datadog
- New Relic
- Honeycomb
Best practice in production:
- services export to OpenTelemetry Collector
- collector batches, filters, samples, and forwards traces
4) Set up sampling carefully
In production, you usually can’t record every trace.
Common options:
- Head-based sampling: decide at trace start
- Tail-based sampling: collector keeps traces and samples based on outcome/latency/errors
For tracing incidents, tail sampling is often better because it can keep:
- slow requests
- errors
- specific routes
- specific users/tenants
5) Correlate logs and metrics with trace IDs
Add trace/span IDs to logs so you can jump from a log line to the trace.
Log fields to include:
trace_idspan_id- service name
- request ID / correlation ID
Also collect metrics like:
- request latency
- error rates
- saturation
- queue lag
6) Make sure async and messaging flows are instrumented
Tracing through microservices often breaks at:
- Kafka
- RabbitMQ
- SQS
- background workers
- async task queues
You need to:
- inject context into message headers/properties
- extract context when consuming
- create producer/consumer spans
7) Use a span structure that makes traces readable
A good pattern:
- root span: incoming API request
- child spans: internal methods
- child spans: outbound HTTP/gRPC calls
- child spans: DB queries or queue publishes
Add useful attributes:
service.namehttp.methodhttp.routehttp.status_codedb.systempeer.service- business fields like
tenant_id,order_idif safe and allowed
8) Verify propagation in production
Check that:
- the same
trace_idappears across services - spans have correct parent-child relationships
- headers are forwarded on all network hops
- sidecars / gateways / proxies are not stripping headers
- CORS / reverse proxies aren’t blocking them
9) If you use Kubernetes / service mesh
OpenTelemetry works well with:
- Kubernetes
- Istio / Linkerd / Envoy
You can instrument apps directly and optionally add:
- Envoy/OpenTelemetry integration
- collector as a DaemonSet or sidecar
- metadata enrichment with pod, namespace, node, cluster
10) Minimal implementation pattern
At a high level:
- initialize tracer provider
- configure exporter
- instrument inbound requests
- inject/extract context on outbound requests
- send to collector/backend
- view trace in UI
Practical example
If Service A receives GET /checkout:
- A starts span
checkout - A calls Service B
/inventory - A injects trace headers
- B extracts them and creates span
inventory-check - B calls Service C
/pricing - same trace continues
- in your tracing UI, you see one trace with all spans in order
Recommended production setup
A strong default architecture:
- app services instrumented with OpenTelemetry SDK/auto-instrumentation
- OpenTelemetry Collector deployed centrally
- tail sampling enabled in collector
- traces + logs + metrics all correlated
- backend: Tempo/Jaeger/Honeycomb/etc.
Common pitfalls
- forgetting to instrument one service
- not propagating headers on outbound calls
- using custom headers instead of standard propagation
- sampling too aggressively
- not instrumenting async workers
- losing trace context at gateways/proxies
- over-instrumenting and adding too much overhead
If you want, I can give you:
- a language-specific setup (Java, Go, Python, Node.js, .NET), or
- a production architecture diagram/pattern for OpenTelemetry tracing across microservices.