Prompt
I'm building an agent workflow, how should I trace every step and model call?
Latest observation
To trace every step and model call in an agent workflow, the key is to treat the workflow like a distributed system: every action should emit a structured event with a shared trace context.
1) Use a trace/span model
Adopt an OpenTelemetry-style structure:
- Trace: one end-to-end agent run
- Span: each step inside the run
- user request received
- planner/model call
- tool invocation
- retrieval/search
- memory read/write
- final response generation
Each span should have:
trace_idspan_idparent_span_id- timestamps
- status: success/error/cancelled
- inputs/outputs references
2) Assign a run ID to every user request
When a request enters the system:
- generate a unique
trace_id - store it in the request context
- propagate it through all downstream calls
If the workflow fans out or has sub-agents, create child spans but keep the same trace.
3) Log structured events, not just text
For every step, emit JSON logs or events like:
{
"trace_id": "abc123",
"span_id": "def456",
"parent_span_id": "root001",
"event_type": "model_call",
"component": "planner",
"model": "gpt-4.1",
"start_time": "2026-10-09T12:00:00Z",
"end_time": "2026-10-09T12:00:01Z",
"input_tokens": 1200,
"output_tokens": 340,
"status": "success"
}
Avoid free-form logs as the primary source of traceability.
4) Instrument every boundary
Trace these points explicitly:
- request ingestion
- prompt construction
- model call start/end
- tool call start/end
- retries
- validation/guardrail checks
- memory or DB reads/writes
- user-visible response generation
If a step is skipped or short-circuited, log that too.
5) Capture model-call metadata
For every model call, record:
- model name/version
- provider
- temperature/top_p
- max tokens
- prompt template ID/version
- input/output token counts
- latency
- finish reason
- error details if any
- cached vs live response
If possible, store prompt and response references, or hashes if you need to reduce sensitive logging.
6) Correlate tool calls with the model that triggered them
A tool call should include:
- which model call decided it
- tool name
- arguments
- result summary
- duration
- retries/error
This makes agent behavior explainable later.
7) Use event sourcing if you need full replay
If you want to reconstruct a run exactly:
- persist every event in order
- store deterministic inputs and outputs
- version prompts, tools, and policies
- record random seeds where applicable
That allows:
- replay
- debugging
- regression testing
- audit trails
8) Store observability data separately from business data
A common pattern:
- traces/logs in an observability backend
- artifacts in blob/object storage
- business records in app DB
Use pointers/IDs between them.
9) Add redaction and privacy controls
Since agent traces can contain sensitive prompts and tool outputs:
- redact secrets/PII before storage
- hash or tokenize sensitive fields
- apply role-based access controls
- define retention policies
10) Recommended minimum schema
At minimum, each event/span should include:
trace_idspan_idparent_span_idstep_namestep_type(model_call,tool_call,retrieval,validation, etc.)timestamp_starttimestamp_endstatusinput_refoutput_referrormetadatamap
11) Practical implementation pattern
A simple approach:
- create a
TraceContextobject - pass it through all functions
- wrap model/tool calls in decorators/middleware
- emit spans automatically on enter/exit
- send everything to OpenTelemetry, LangSmith, or your own collector
12) Good defaults
If you’re starting from scratch:
- use OpenTelemetry for trace IDs/spans
- use structured JSON logs
- keep prompt/tool/version metadata
- add redaction
- make every agent step a span
If you want, I can also give you:
- a sample event schema,
- Python/TypeScript instrumentation code, or
- an OpenTelemetry design for agent workflows.
Brands