Prompt
How should I set up observability for a multi-step agent so I can see every model call and tool use?
Latest observation
To observe a multi-step agent well, design telemetry around each hop in the loop:
- Agent step / turn
- Model call
- Tool call
- Tool result / error
- Final response
The goal is to make every action traceable with a shared trace ID and nested span IDs.
Recommended setup
1) Use distributed tracing
Instrument the agent with traces/spans so every run becomes a tree:
- Trace = one user request / one agent run
- Span = one model invocation, one tool call, one retrieval step, etc.
Typical span hierarchy:
agent.runllm.calltool.searchllm.calltool.database_queryllm.callagent.finalize
Attach:
trace_idspan_idparent_span_id- timestamps
- duration
- status (
ok,error,retry) - input/output summaries
- token usage
- tool args/results
2) Log structured events for each transition
In addition to traces, emit JSON logs for:
- prompt sent to model
- model output
- tool selected
- tool arguments
- tool output
- exceptions
- retries / fallbacks
Use a consistent schema so logs can be queried easily.
Example event types:
agent_startedllm_requestllm_responsetool_invokedtool_returnedtool_failedagent_finished
3) Capture model-level telemetry
For every model call, record:
- model name/version
- prompt template or prompt hash
- full prompt or redacted prompt
- response text or redacted response
- tool/function calls requested by the model
- latency
- token counts
- cost estimate
- finish reason
- safety/filtering flags
If the model supports tool/function calling, log:
- requested tool name
- arguments
- parsed arguments
- validation errors
- whether the call was executed
4) Capture tool-level telemetry
For each tool use, record:
- tool name
- normalized input arguments
- execution time
- output size
- output preview
- retry attempts
- exception details
- upstream dependency status
If tools access external systems, also instrument those systems if possible.
5) Correlate everything with request context
Propagate metadata through the full agent execution:
user_idsession_idconversation_idrequest_id- environment (
dev,staging,prod) - agent version / prompt version
- feature flags
- tenant/org ID
6) Redact secrets and sensitive data
Since observability often captures prompts and tool inputs/outputs, add:
- PII scrubbing
- secret detection
- allowlist-based field capture
- sampling in production
- payload truncation
Practical implementation pattern
If your agent has a loop like:
- send messages to model
- receive plan or tool call
- execute tool
- append result
- call model again
Wrap each iteration with spans:
trace: agent.run
span: llm.call #1
span: tool.search
span: llm.call #2
span: tool.fetch_document
span: llm.call #3
And for each span, emit an event or structured log.
Best tools to use
A common stack is:
- OpenTelemetry for traces/metrics/log correlation
- OpenTelemetry Collector to route data
- Backend such as:
- Jaeger / Tempo for traces
- Grafana / Prometheus for metrics
- ELK / Loki / Datadog / New Relic for logs and dashboards
If you use an agent framework, check whether it already supports tracing:
- LangChain / LangGraph callbacks
- OpenAI / Anthropic SDK wrappers
- custom middleware in your orchestration layer
Useful metrics
Track these per agent version and per task type:
- success rate
- tool call count per run
- average model calls per run
- average latency per step
- tool failure rate
- retry rate
- token usage and cost
- abandonment rate
- hallucination / validation failure rate
- time to first useful action
Good dashboard views
Build dashboards for:
- single run replay: see exact sequence of model/tool calls
- timeline: latency per step
- error hotspots: failing tools / prompt versions
- cost breakdown: tokens and external tool costs
- tool effectiveness: which tools actually help
- prompt version comparison: old vs new prompt outcomes
Minimal schema example
{
"trace_id": "abc123",
"span_id": "s1",
"parent_span_id": null,
"type": "llm.call",
"agent_name": "support_agent",
"model": "gpt-4.1",
"input": {
"messages_count": 6,
"prompt_hash": "p789"
},
"output": {
"finish_reason": "tool_call",
"tool_name": "search_kb"
},
"metrics": {
"latency_ms": 842,
"prompt_tokens": 1234,
"completion_tokens": 98,
"cost_usd": 0.012
},
"status": "ok",
"timestamp": "2026-10-09T12:00:00Z"
}
If you want the simplest effective setup
Start with:
- OpenTelemetry traces
- JSON logs
- A wrapper around every model/tool call
- Correlation IDs propagated through the agent loop
That’s enough to replay and debug nearly any multi-step agent.
If you want, I can also give you:
- a reference architecture,
- a Python example using OpenTelemetry,
- or a LangChain/LangGraph-specific setup.