Prompt
I need tracing for every step in a chain of LLM calls
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
To trace every step in a chain of LLM calls, you typically want distributed tracing plus structured logs. The key idea is to assign a single trace ID to the whole request and a span ID to each LLM call and sub-step.
What to capture
For each step in the chain, record:
- trace_id: same across the entire end-to-end workflow
- span_id: unique per step / LLM call
- parent_span_id: links the step to the previous step
- step name: e.g.
planner,retriever,tool_call,summarizer - model used
- prompt / input
- response / output
- token counts
- latency
- errors / retries / timeout
- tool calls and tool outputs, if any
Recommended structure
A trace should look like a tree or chain:
request
├─ planner LLM call
├─ retriever step
│ └─ embedding / search LLM call
├─ tool execution
└─ final answer LLM call
Minimal JSON example
{
"trace_id": "trace-123",
"span_id": "span-1",
"parent_span_id": null,
"step": "user_request",
"timestamp": "2026-10-09T12:00:00Z",
"input": "What are the risks?",
"output": null
}
{
"trace_id": "trace-123",
"span_id": "span-2",
"parent_span_id": "span-1",
"step": "planner_llm",
"model": "gpt-4.1",
"prompt": "Plan the analysis",
"response": "1. Retrieve docs 2. Summarize risks",
"latency_ms": 842,
"tokens_in": 120,
"tokens_out": 48
}
Best practice
Use an observability standard like OpenTelemetry:
- create a root trace when the request starts
- create a child span for each LLM call
- propagate context through all internal functions and tool calls
- export spans to a backend such as:
- OpenTelemetry Collector
- Jaeger
- Grafana Tempo
- Datadog
- Honeycomb
If you’re using an LLM framework
Many frameworks already support tracing:
- LangChain → LangSmith / callbacks
- LlamaIndex → observability hooks
- OpenAI-style apps → wrap each API call in an OpenTelemetry span
Practical implementation pattern
- Generate
trace_idat request start - Pass
trace_idand currentspan_idthrough all functions - Wrap each LLM call in a span
- Log prompts, outputs, and metadata
- Export traces to your observability system
Example pseudocode
with tracer.start_as_current_span("root_request") as root:
root.set_attribute("trace_id", trace_id)
with tracer.start_as_current_span("planner_llm") as span:
span.set_attribute("model", "gpt-4.1")
span.set_attribute("prompt", prompt)
response = call_llm(prompt)
span.set_attribute("response", response)
with tracer.start_as_current_span("tool_search") as span:
results = search_docs(query)
with tracer.start_as_current_span("final_llm") as span:
answer = call_llm(build_prompt(results))
If you want, I can also give you:
- a Python OpenTelemetry example
- a LangChain tracing setup
- a production-ready JSON schema for chain tracing