Prompt
How do I trace a multi-step agent across prompt, retrieval, and tool calls?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
To trace a multi-step agent end to end, treat each run as a single trace with nested spans for each phase:
- Prompt construction
- Retrieval / context fetching
- LLM reasoning / planning
- Tool execution
- Final response
What to log
For each step, capture:
- trace_id: one ID for the whole agent run
- span_id: one ID per step
- parent_span_id: to build the tree
- timestamps: start/end
- inputs/outputs: prompt, retrieved docs, tool args/results
- metadata: model name, tool name, retrieval query, token counts, latency, errors
Typical span structure
Example:
trace: agent_run_123span: prompt_buildspan: retrievalspan: vector_searchspan: rerank
span: llm_callspan: tool_call:weather_apispan: llm_call_2span: final_answer
How to implement
1) Wrap each stage
Use a tracing library or OpenTelemetry-style spans.
Pseudo-code:
with trace("agent_run") as trace:
with trace.span("prompt_build") as s:
prompt = build_prompt(user_input)
with trace.span("retrieval") as s:
docs = retriever.search(user_input)
with trace.span("llm_call") as s:
plan = llm(prompt + docs)
for tool in plan.tools:
with trace.span(f"tool:{tool.name}") as s:
result = tool.execute(tool.args)
with trace.span("final_llm") as s:
answer = llm(make_final_prompt(plan, results))
2) Preserve context across steps
Pass the same trace context through:
- prompt builder
- retriever
- tool executor
- agent planner
- post-processor
If using async/concurrent execution, propagate the trace context explicitly.
3) Record artifacts separately
Large payloads can be expensive to store inline. Common practice:
- store prompt text, retrieved chunks, and tool I/O
- or store references/IDs and keep full content in object storage
4) Correlate retrieval with generation
Make sure retrieval span records:
- query text
- filters
- top-k results
- doc IDs / chunk IDs
- scores
- rerank decisions
Then attach those doc IDs to the prompt span so you can see exactly which context influenced the model.
5) Correlate tool calls with decisions
For tool spans, log:
- tool name/version
- arguments
- returned value
- errors/retries
- whether the output was used in the next LLM call
Best practices
- Use a single trace ID per user request or agent task.
- Use nested spans for every internal step.
- Redact secrets and PII before logging.
- Log structured data, not just free-form text.
- Capture retry attempts as separate spans or events.
- Include token usage and latency for each LLM call.
- If the agent branches, represent each branch as sibling spans.
Example trace
agent_run (trace_id=abc)
├── prompt_build
├── retrieval
│ ├── vector_search
│ └── rerank
├── llm_plan
├── tool_call:calendar.lookup
├── tool_call:email.send
└── llm_finalize
If you want observability tooling
Common options:
- OpenTelemetry
- LangSmith
- Arize Phoenix
- Helicone
- Weights & Biases Traces
If you tell me your stack (Python/JS, LangChain/LlamaIndex/custom, OpenTelemetry or not), I can show a concrete tracing pattern or code sample.