Prompt
What should I use for observability on a multi-step agent in production?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
For a multi-step agent in production, you’ll usually want three layers of observability, not just one tool:
1) Tracing for step-by-step execution
Use distributed tracing to see each agent step, tool call, LLM call, retrieval, retry, and latency breakdown.
Good options:
- OpenTelemetry (best foundation; vendor-neutral)
- A tracing backend like Datadog, Honeycomb, Grafana Tempo, Jaeger, or New Relic
- If you’re using an agent framework, many support trace hooks that can export to OTel
What to capture:
- Trace per user request / conversation turn
- Spans for:
- model calls
- tool invocations
- retrieval
- memory/state reads/writes
- retries/timeouts
- guardrail checks
- Span attributes:
- model name/version
- prompt template version
- tool name
- token counts
- latency
- success/failure status
- cost estimate
- user/session/tenant IDs (careful with privacy)
2) Structured logs for debugging and audits
Use structured JSON logging for:
- prompts and outputs (with redaction)
- tool inputs/outputs
- errors/exceptions
- agent decisions
- state transitions
This is what you’ll use when traces show where something went wrong, but you need to know why.
Important:
- redact secrets, PII, and internal system prompts if needed
- keep correlation IDs so logs tie back to traces
3) Evaluation/quality monitoring for “is the agent good?”
Observability isn’t just runtime health — you also need behavioral monitoring:
- task success rate
- tool-call accuracy
- hallucination / unsupported claim rate
- latency to completion
- token/cost per task
- user satisfaction / thumbs up-down
- escalation rate to human
- retrieval precision/recall (if RAG)
- regression tests on golden conversations
Useful tools:
- LangSmith (great for LLM/agent tracing + evals)
- Weights & Biases Weave
- Arize Phoenix
- Humanloop
- Helicone (LLM request logging/analytics)
- PromptLayer (prompt/version tracking)
- OpenAI Evals or custom eval pipelines
A practical production stack
If you want a solid default:
- OpenTelemetry for traces
- Structured logging to ELK / Datadog / CloudWatch / Stackdriver
- Prometheus + Grafana for metrics
- LangSmith / Phoenix / Weave for LLM-specific tracing and evals
What matters most for multi-step agents
Multi-step agents fail in ways normal apps don’t, so make sure you can answer:
- What did the agent do at each step?
- Which tool/model call caused the issue?
- Was the failure due to retrieval, reasoning, tool use, or state?
- Did the agent complete the task correctly?
- How much did it cost and how long did it take?
If you want the shortest recommendation
- Best general-purpose foundation: OpenTelemetry
- Best LLM/agent-specific observability: LangSmith or Arize Phoenix
- Best combo in production: OTel + structured logs + eval pipeline
If you tell me your stack (LangChain/LlamaIndex/custom, cloud provider, and whether you need compliance or on-prem), I can recommend a more specific setup.