Prompt
What should I use for tracing tool calls in agent workflows?
Latest observation
For tracing tool calls in agent workflows, you generally want distributed tracing + structured event logs with spans around each agent step and tool invocation.
Best practice
Use OpenTelemetry (OTel) as the backbone:
- One trace per user request / agent run
- Nested spans for:
- model calls
- planning steps
- tool calls
- retries / fallbacks
- memory / retrieval lookups
- Add structured attributes like:
tool.nametool.inputtool.outputagent.stepmodel.nametoken_usagelatency_mserror.type
Good options to pair with it
Depending on how much LLM-specific visibility you want:
-
OpenTelemetry + Grafana/Tempo/Jaeger
Best general-purpose tracing stack. -
LangSmith
Great if you’re using LangChain/LangGraph and want agent/tool traces with minimal setup. -
Arize Phoenix
Strong for LLM observability, evals, retrieval tracing, and debugging agent behavior. -
Honeycomb / Datadog / New Relic
Good if you already use one of these for production observability.
What to capture for tool calls
At minimum:
- tool name
- start/end timestamps
- input parameters
- output/result
- success/failure
- error message/stack
- retries
- correlation IDs / parent span ID
Recommendation
If you want the most future-proof approach:
- Instrument everything with OpenTelemetry
- Export traces to Grafana Tempo, Jaeger, or your observability vendor
- If you need LLM-specific debugging, layer in LangSmith or Phoenix
If you want, I can also suggest a concrete setup for Python, JS, or LangGraph.