Prompt

What should I use for OpenTelemetry-style tracing in LLM systems?

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

For OpenTelemetry-style tracing in LLM systems, the best fit is usually OpenTelemetry itself, with an LLM-specific instrumentation layer on top.

Recommended approach

  • Use OpenTelemetry (OTel) for the core tracing model:

    • spans
    • trace/context propagation
    • attributes/metadata
    • exporters to Jaeger, Tempo, Honeycomb, Datadog, etc.
  • Add LLM-aware conventions/instrumentation to capture things OTel doesn’t define natively, such as:

    • prompt / completion metadata
    • model name and version
    • token counts
    • latency per generation step
    • tool calls / function calls
    • retrieval spans
    • agent planning / memory steps
    • safety/filtering events

Common practical options

  1. OpenTelemetry + custom spans/attributes
    Best if you want vendor-neutral observability and already use OTel.

  2. LangSmith Best if you’re building with LangChain and want turnkey LLM tracing/debugging.

  3. Arize Phoenix Good for tracing plus evals and debugging RAG/agent systems.

  4. Helicone Useful for LLM request tracing, cost, latency, and API observability.

  5. OpenLIT / OpenInference Helpful if you want more standardized LLM tracing with OpenTelemetry compatibility.

What I’d recommend

If your goal is specifically “OpenTelemetry-style”, use:

  • OpenTelemetry SDK
  • OpenInference semantic conventions or your own consistent span schema
  • auto-instrumentation wrappers around LLM calls, retrieval, and tool use
  • an OTel backend like Jaeger, Tempo, Honeycomb, Grafana, Datadog, or New Relic

Good span structure for LLM apps

Typical trace:

  • request
    • retrieval
      • vector_search
      • rerank
    • prompt निर्माण / assembly
    • llm.chat_completion
    • tool.call
    • postprocess

Key fields to record

  • llm.model_name
  • llm.provider
  • llm.temperature
  • llm.max_tokens
  • llm.input_tokens
  • llm.output_tokens
  • llm.total_tokens
  • llm.system_prompt / llm.prompt_template_id
  • rag.top_k
  • tool.name
  • error.type, error.message

One caution

Be careful not to log raw prompts or outputs if they may contain sensitive data. Prefer:

  • hashes
  • redaction
  • sampling
  • structured metadata
  • secure storage for full text only when needed

If you want, I can also give you:

  1. a recommended OTel span schema for LLM apps, or
  2. Python code showing OpenTelemetry tracing around an LLM call.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.