Prompt
I'm building an agent that calls tools, how do I record all the intermediate steps?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
To record all intermediate steps in a tool-calling agent, the usual pattern is:
-
Log every model action
- Prompt sent to the model
- Model output
- Any tool call name + arguments
-
Log every tool execution
- Tool name
- Inputs
- Outputs
- Errors / retries
- Timestamps / latency
-
Store the full trace as a sequence of events A simple structure is:
[
{"type": "llm_request", "prompt": "..."},
{"type": "llm_response", "content": "...", "tool_call": {"name": "search", "args": {"q": "weather"}}},
{"type": "tool_request", "name": "search", "args": {"q": "weather"}},
{"type": "tool_response", "name": "search", "result": "..."},
{"type": "llm_response", "content": "It is sunny."}
]
-
If you’re using an agent framework, enable tracing/callbacks Most frameworks expose hooks like:
on_llm_start,on_llm_endon_tool_start,on_tool_endon_chain_start,on_chain_end
-
Keep a run/session ID This lets you connect all intermediate steps from one agent invocation.
-
Persist to a datastore or observability system Common options:
- SQLite/Postgres for simple persistence
- JSONL files for easy debugging
- OpenTelemetry / tracing backends for production
- Framework-specific tracing tools if available
Example pseudo-code:
trace = []
def log(event):
trace.append({**event, "ts": time.time()})
log({"type": "llm_request", "messages": messages})
resp = llm(messages)
log({"type": "llm_response", "output": resp})
if resp.tool_call:
log({"type": "tool_request", "tool": resp.tool_call.name, "args": resp.tool_call.args})
tool_result = tools[resp.tool_call.name](**resp.tool_call.args)
log({"type": "tool_response", "result": tool_result})
If you want, I can also show:
- how to do this in LangChain
- how to do this in OpenAI tool calling
- or a framework-agnostic event schema for agent traces