Prompt
What should I use for OpenTelemetry-style tracing in LLM systems?
Latest observation
For OpenTelemetry-style tracing in LLM systems, the best fit is usually OpenTelemetry itself, with an LLM-specific instrumentation layer on top.
Recommended approach
-
Use OpenTelemetry (OTel) for the core tracing model:
- spans
- trace/context propagation
- attributes/metadata
- exporters to Jaeger, Tempo, Honeycomb, Datadog, etc.
-
Add LLM-aware conventions/instrumentation to capture things OTel doesn’t define natively, such as:
- prompt / completion metadata
- model name and version
- token counts
- latency per generation step
- tool calls / function calls
- retrieval spans
- agent planning / memory steps
- safety/filtering events
Common practical options
-
OpenTelemetry + custom spans/attributes
Best if you want vendor-neutral observability and already use OTel. -
LangSmith Best if you’re building with LangChain and want turnkey LLM tracing/debugging.
-
Arize Phoenix Good for tracing plus evals and debugging RAG/agent systems.
-
Helicone Useful for LLM request tracing, cost, latency, and API observability.
-
OpenLIT / OpenInference Helpful if you want more standardized LLM tracing with OpenTelemetry compatibility.
What I’d recommend
If your goal is specifically “OpenTelemetry-style”, use:
- OpenTelemetry SDK
- OpenInference semantic conventions or your own consistent span schema
- auto-instrumentation wrappers around LLM calls, retrieval, and tool use
- an OTel backend like Jaeger, Tempo, Honeycomb, Grafana, Datadog, or New Relic
Good span structure for LLM apps
Typical trace:
requestretrievalvector_searchrerank
prompt निर्माण / assemblyllm.chat_completiontool.callpostprocess
Key fields to record
llm.model_namellm.providerllm.temperaturellm.max_tokensllm.input_tokensllm.output_tokensllm.total_tokensllm.system_prompt/llm.prompt_template_idrag.top_ktool.nameerror.type,error.message
One caution
Be careful not to log raw prompts or outputs if they may contain sensitive data. Prefer:
- hashes
- redaction
- sampling
- structured metadata
- secure storage for full text only when needed
If you want, I can also give you:
- a recommended OTel span schema for LLM apps, or
- Python code showing OpenTelemetry tracing around an LLM call.