Prompt

I'm building an agent workflow, how should I trace every step and model call?

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

To trace every step and model call in an agent workflow, the key is to treat the workflow like a distributed system: every action should emit a structured event with a shared trace context.

1) Use a trace/span model

Adopt an OpenTelemetry-style structure:

  • Trace: one end-to-end agent run
  • Span: each step inside the run
    • user request received
    • planner/model call
    • tool invocation
    • retrieval/search
    • memory read/write
    • final response generation

Each span should have:

  • trace_id
  • span_id
  • parent_span_id
  • timestamps
  • status: success/error/cancelled
  • inputs/outputs references

2) Assign a run ID to every user request

When a request enters the system:

  • generate a unique trace_id
  • store it in the request context
  • propagate it through all downstream calls

If the workflow fans out or has sub-agents, create child spans but keep the same trace.

3) Log structured events, not just text

For every step, emit JSON logs or events like:

{
  "trace_id": "abc123",
  "span_id": "def456",
  "parent_span_id": "root001",
  "event_type": "model_call",
  "component": "planner",
  "model": "gpt-4.1",
  "start_time": "2026-10-09T12:00:00Z",
  "end_time": "2026-10-09T12:00:01Z",
  "input_tokens": 1200,
  "output_tokens": 340,
  "status": "success"
}

Avoid free-form logs as the primary source of traceability.

4) Instrument every boundary

Trace these points explicitly:

  • request ingestion
  • prompt construction
  • model call start/end
  • tool call start/end
  • retries
  • validation/guardrail checks
  • memory or DB reads/writes
  • user-visible response generation

If a step is skipped or short-circuited, log that too.

5) Capture model-call metadata

For every model call, record:

  • model name/version
  • provider
  • temperature/top_p
  • max tokens
  • prompt template ID/version
  • input/output token counts
  • latency
  • finish reason
  • error details if any
  • cached vs live response

If possible, store prompt and response references, or hashes if you need to reduce sensitive logging.

6) Correlate tool calls with the model that triggered them

A tool call should include:

  • which model call decided it
  • tool name
  • arguments
  • result summary
  • duration
  • retries/error

This makes agent behavior explainable later.

7) Use event sourcing if you need full replay

If you want to reconstruct a run exactly:

  • persist every event in order
  • store deterministic inputs and outputs
  • version prompts, tools, and policies
  • record random seeds where applicable

That allows:

  • replay
  • debugging
  • regression testing
  • audit trails

8) Store observability data separately from business data

A common pattern:

  • traces/logs in an observability backend
  • artifacts in blob/object storage
  • business records in app DB

Use pointers/IDs between them.

9) Add redaction and privacy controls

Since agent traces can contain sensitive prompts and tool outputs:

  • redact secrets/PII before storage
  • hash or tokenize sensitive fields
  • apply role-based access controls
  • define retention policies

10) Recommended minimum schema

At minimum, each event/span should include:

  • trace_id
  • span_id
  • parent_span_id
  • step_name
  • step_type (model_call, tool_call, retrieval, validation, etc.)
  • timestamp_start
  • timestamp_end
  • status
  • input_ref
  • output_ref
  • error
  • metadata map

11) Practical implementation pattern

A simple approach:

  • create a TraceContext object
  • pass it through all functions
  • wrap model/tool calls in decorators/middleware
  • emit spans automatically on enter/exit
  • send everything to OpenTelemetry, LangSmith, or your own collector

12) Good defaults

If you’re starting from scratch:

  • use OpenTelemetry for trace IDs/spans
  • use structured JSON logs
  • keep prompt/tool/version metadata
  • add redaction
  • make every agent step a span

If you want, I can also give you:

  1. a sample event schema,
  2. Python/TypeScript instrumentation code, or
  3. an OpenTelemetry design for agent workflows.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.