Prompt

How should I set up observability for a multi-step agent so I can see every model call and tool use?

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

To observe a multi-step agent well, design telemetry around each hop in the loop:

  1. Agent step / turn
  2. Model call
  3. Tool call
  4. Tool result / error
  5. Final response

The goal is to make every action traceable with a shared trace ID and nested span IDs.

Recommended setup

1) Use distributed tracing

Instrument the agent with traces/spans so every run becomes a tree:

  • Trace = one user request / one agent run
  • Span = one model invocation, one tool call, one retrieval step, etc.

Typical span hierarchy:

  • agent.run
    • llm.call
    • tool.search
    • llm.call
    • tool.database_query
    • llm.call
    • agent.finalize

Attach:

  • trace_id
  • span_id
  • parent_span_id
  • timestamps
  • duration
  • status (ok, error, retry)
  • input/output summaries
  • token usage
  • tool args/results

2) Log structured events for each transition

In addition to traces, emit JSON logs for:

  • prompt sent to model
  • model output
  • tool selected
  • tool arguments
  • tool output
  • exceptions
  • retries / fallbacks

Use a consistent schema so logs can be queried easily.

Example event types:

  • agent_started
  • llm_request
  • llm_response
  • tool_invoked
  • tool_returned
  • tool_failed
  • agent_finished

3) Capture model-level telemetry

For every model call, record:

  • model name/version
  • prompt template or prompt hash
  • full prompt or redacted prompt
  • response text or redacted response
  • tool/function calls requested by the model
  • latency
  • token counts
  • cost estimate
  • finish reason
  • safety/filtering flags

If the model supports tool/function calling, log:

  • requested tool name
  • arguments
  • parsed arguments
  • validation errors
  • whether the call was executed

4) Capture tool-level telemetry

For each tool use, record:

  • tool name
  • normalized input arguments
  • execution time
  • output size
  • output preview
  • retry attempts
  • exception details
  • upstream dependency status

If tools access external systems, also instrument those systems if possible.

5) Correlate everything with request context

Propagate metadata through the full agent execution:

  • user_id
  • session_id
  • conversation_id
  • request_id
  • environment (dev, staging, prod)
  • agent version / prompt version
  • feature flags
  • tenant/org ID

6) Redact secrets and sensitive data

Since observability often captures prompts and tool inputs/outputs, add:

  • PII scrubbing
  • secret detection
  • allowlist-based field capture
  • sampling in production
  • payload truncation

Practical implementation pattern

If your agent has a loop like:

  1. send messages to model
  2. receive plan or tool call
  3. execute tool
  4. append result
  5. call model again

Wrap each iteration with spans:

trace: agent.run
  span: llm.call #1
  span: tool.search
  span: llm.call #2
  span: tool.fetch_document
  span: llm.call #3

And for each span, emit an event or structured log.

Best tools to use

A common stack is:

  • OpenTelemetry for traces/metrics/log correlation
  • OpenTelemetry Collector to route data
  • Backend such as:
    • Jaeger / Tempo for traces
    • Grafana / Prometheus for metrics
    • ELK / Loki / Datadog / New Relic for logs and dashboards

If you use an agent framework, check whether it already supports tracing:

  • LangChain / LangGraph callbacks
  • OpenAI / Anthropic SDK wrappers
  • custom middleware in your orchestration layer

Useful metrics

Track these per agent version and per task type:

  • success rate
  • tool call count per run
  • average model calls per run
  • average latency per step
  • tool failure rate
  • retry rate
  • token usage and cost
  • abandonment rate
  • hallucination / validation failure rate
  • time to first useful action

Good dashboard views

Build dashboards for:

  • single run replay: see exact sequence of model/tool calls
  • timeline: latency per step
  • error hotspots: failing tools / prompt versions
  • cost breakdown: tokens and external tool costs
  • tool effectiveness: which tools actually help
  • prompt version comparison: old vs new prompt outcomes

Minimal schema example

{
  "trace_id": "abc123",
  "span_id": "s1",
  "parent_span_id": null,
  "type": "llm.call",
  "agent_name": "support_agent",
  "model": "gpt-4.1",
  "input": {
    "messages_count": 6,
    "prompt_hash": "p789"
  },
  "output": {
    "finish_reason": "tool_call",
    "tool_name": "search_kb"
  },
  "metrics": {
    "latency_ms": 842,
    "prompt_tokens": 1234,
    "completion_tokens": 98,
    "cost_usd": 0.012
  },
  "status": "ok",
  "timestamp": "2026-10-09T12:00:00Z"
}

If you want the simplest effective setup

Start with:

  1. OpenTelemetry traces
  2. JSON logs
  3. A wrapper around every model/tool call
  4. Correlation IDs propagated through the agent loop

That’s enough to replay and debug nearly any multi-step agent.

If you want, I can also give you:

  • a reference architecture,
  • a Python example using OpenTelemetry,
  • or a LangChain/LangGraph-specific setup.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.