Prompt
How do I instrument microservices for tracing with OpenTelemetry
Latest observation
To instrument microservices for tracing with OpenTelemetry, you generally want to do 5 things:
- Choose your tracing backend
- Add OpenTelemetry SDK/auto-instrumentation to each service
- Propagate context between services
- Export traces to an OpenTelemetry Collector or backend
- Verify spans, services, and trace continuity
1) Pick the architecture
A common setup is:
- Microservices emit spans via OpenTelemetry SDK
- They send traces to an OpenTelemetry Collector
- The collector exports to Jaeger, Grafana Tempo, Datadog, Honeycomb, Zipkin, etc.
This is preferred because the collector handles batching, retries, sampling, and backend portability.
2) Instrument each service
Option A: Auto-instrumentation
This is the fastest way, especially for common frameworks.
- Java: OpenTelemetry Java agent
- Python:
opentelemetry-instrument - Node.js: OpenTelemetry auto-instrumentation packages
- .NET: OpenTelemetry .NET instrumentation packages
- Go: usually manual instrumentation, though some auto options exist depending on stack
Auto-instrumentation captures incoming HTTP/gRPC requests, outgoing calls, DB queries, etc.
Option B: Manual instrumentation
Use this when you need custom spans or your framework isn’t fully covered.
Add:
- a server span when a request enters the service
- client spans for outbound calls
- custom spans around important business logic
- span attributes for metadata like order ID, tenant, region, etc.
3) Propagate trace context across services
This is the key part for microservices.
Use standard propagation:
- W3C Trace Context (
traceparent,tracestate) - optionally Baggage
When service A calls service B:
- A injects trace context into HTTP/gRPC headers
- B extracts context and continues the same trace
If context propagation is missing, each service creates separate traces and you lose end-to-end visibility.
4) Export traces
Configure:
- Tracer provider
- Resource attributes like service name, environment, version
- OTLP exporter to the collector
Typical resource attributes:
service.nameservice.namespaceservice.versiondeployment.environment
OTLP is the standard export protocol.
5) Add sampling and performance controls
In microservices, tracing overhead matters.
Use:
- head-based sampling for simple control
- tail-based sampling in the collector if you want to keep slow/error traces
- batching/export queue settings
- ignore health checks and noisy endpoints
A good collector setup can do intelligent sampling centrally.
Example: Python
from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
resource = Resource.create({
"service.name": "orders-service",
"service.version": "1.0.0",
"deployment.environment": "prod"
})
provider = TracerProvider(resource=resource)
trace.set_tracer_provider(provider)
exporter = OTLPSpanExporter(endpoint="http://otel-collector:4317", insecure=True)
processor = BatchSpanProcessor(exporter)
provider.add_span_processor(processor)
tracer = trace.get_tracer(__name__)
with tracer.start_as_current_span("process_order"):
# business logic
pass
If you use a web framework, also add HTTP instrumentation.
Example: Node.js
const { NodeSDK } = require('@opentelemetry/sdk-node');
const { OTLPTraceExporter } = require('@opentelemetry/exporter-trace-otlp-grpc');
const { Resource } = require('@opentelemetry/resources');
const { SemanticResourceAttributes } = require('@opentelemetry/semantic-conventions');
const sdk = new NodeSDK({
resource: new Resource({
[SemanticResourceAttributes.SERVICE_NAME]: 'orders-service',
}),
traceExporter: new OTLPTraceExporter({
url: 'http://otel-collector:4317',
}),
});
sdk.start();
Example: Java
Use the OpenTelemetry Java agent:
java -javaagent:/path/opentelemetry-javaagent.jar \
-Dotel.service.name=orders-service \
-Dotel.exporter.otlp.endpoint=http://otel-collector:4317 \
-jar app.jar
This is often the easiest option for Java microservices.
Best practices
- Use consistent service names
- Trace only important requests at first
- Add custom spans for:
- payment authorization
- inventory checks
- third-party API calls
- async jobs/queues
- Include useful span attributes, but avoid high-cardinality values unless necessary
- Make sure background jobs and message consumers continue traces too
- Instrument queues/events as well as HTTP/gRPC
- Add logs with trace IDs for correlation
Common pitfalls
- Forgetting context propagation
- Using different service names per instance instead of per logical service
- Tracing everything with no sampling
- Adding high-cardinality attributes like full user IDs or raw URLs
- Not instrumenting async/message-driven flows
- Exporting directly from every service without a collector in larger systems
A good rollout plan
- Instrument one service end-to-end
- Add outbound call propagation
- Add the next downstream service
- Deploy an OpenTelemetry Collector
- Add sampling and dashboards
- Expand to all services and async workers
If you want, I can also give you:
- a language-specific setup for Java, Python, Node.js, Go, or .NET
- a Docker Compose example with OpenTelemetry Collector + Jaeger/Tempo
- a Kubernetes deployment pattern
- a full microservices tracing example with HTTP calls between services