Prompt
I need observability for an agent that uses multiple tools and APIs
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Absolutely — for an agent that uses multiple tools and APIs, observability should let you answer four core questions:
- What did the agent try to do?
- Which tool/API calls happened, in what order, and with what inputs?
- What happened at each step — latency, errors, retries, token usage, cost?
- Why did the agent make the final decision/output?
What to instrument
1) End-to-end traces
Create a single trace per user request / agent run, with spans for:
- user request received
- model inference step(s)
- tool call invocation
- external API request/response
- post-processing / decision step
- final response generation
Use a correlation ID / trace ID across all components.
2) Structured logs
Log structured events for each step:
run_startedreasoning_step_started/reasoning_step_completedtool_selectedtool_call_startedtool_call_completedtool_call_failedretry_scheduledfinal_answer_generated
Include:
- trace/run ID
- tool name
- input parameters (redacted where needed)
- output summary
- timestamps
- duration
- status/error code
- model name/version
- token counts
- cost estimates
3) Metrics
Track aggregate metrics:
- request latency p50/p95/p99
- tool call success rate
- tool-specific latency
- retry rate
- timeout rate
- error rate by tool/API
- tokens per run
- cost per run
- completion success / task success rate
- hallucinated tool call rate / invalid tool args rate
4) Prompt / reasoning observability
For LLM-based agents, capture:
- prompt version
- system prompt
- tool definitions available at the time
- model configuration
- selected tool and rationale if available
- output schema validation results
If you can’t store full chain-of-thought, store decision summaries or step metadata instead.
5) Tool/API-specific instrumentation
For every external API/tool:
- request endpoint
- request payload shape
- response status
- response time
- retries/backoff
- rate limits hit
- circuit breaker events
- dependency errors
Recommended event schema
A common schema for all agent events helps a lot:
{
"trace_id": "abc123",
"run_id": "run_789",
"span_id": "span_456",
"parent_span_id": "span_123",
"event_type": "tool_call_completed",
"timestamp": "2026-10-09T12:34:56.789Z",
"agent_name": "support-agent",
"step_name": "lookup_customer",
"tool_name": "crm_api",
"model_name": "gpt-4.1",
"status": "success",
"duration_ms": 184,
"input": {
"customer_id": "redacted"
},
"output_summary": {
"records_found": 1
},
"error": null,
"token_usage": {
"prompt": 512,
"completion": 128
}
}
Best practices
- Use distributed tracing with OpenTelemetry if possible.
- Redact secrets and PII before logging.
- Version everything: prompts, tool schemas, policies, model versions.
- Capture failures explicitly: timeouts, malformed args, tool exceptions, model refusals.
- Record retries and fallback paths.
- Measure business outcomes, not just system health.
- Link logs, traces, and metrics via trace IDs.
Practical stack suggestion
A common setup:
- Tracing: OpenTelemetry + Jaeger / Tempo / Datadog
- Logs: JSON logs to ELK / Loki / Datadog
- Metrics: Prometheus / Grafana / Datadog
- LLM/agent-specific observability: Langfuse, Arize Phoenix, Helicone, OpenTelemetry integrations
Minimal architecture
- agent runtime emits spans/events
- tool wrappers emit child spans
- all logs include
trace_id - metrics exported from runtime
- dashboards by agent, tool, version, and failure mode
Dashboards to build
- run success rate over time
- latency breakdown by step
- tool failure heatmap
- retries and timeouts
- cost per successful task
- model version comparison
- prompt version comparison
- top failing tool/API dependencies
If you want, I can also give you:
- a reference architecture,
- a Python/OpenTelemetry implementation example, or
- a logging schema for multi-tool agents.