Prompt
How do I trace prompts, model calls, tool use, and final responses in an LLM app?
Latest observation
To trace prompts, model calls, tool use, and final responses in an LLM app, you want end-to-end observability across the whole request lifecycle. The usual pattern is:
- Assign a unique trace/request ID
- Log every step as a structured event
- Link related events with parent/child spans
- Capture inputs/outputs for prompts, model calls, tools, and final answers
- Send it all to a tracing backend or logging system
What to trace
At minimum, capture these stages:
- User request received
- Prompt construction
- system prompt
- developer prompt
- user prompt
- retrieved context
- Model call
- model name/version
- parameters: temperature, max tokens, top_p, etc.
- request payload
- response payload
- latency
- token usage
- errors/retries
- Tool invocation
- tool name
- arguments
- input/output
- latency
- error
- Post-processing
- parsing
- reranking
- guardrails
- Final response
- answer text
- citations
- confidence or refusal reason if applicable
Best practice: use spans
A good mental model is distributed tracing:
- Trace = one user interaction
- Span = one step inside it
Example span tree:
requestprompt_buildretrievalllm_calltool_call:get_weathertool_call:search_docs
response_formatting
Each span should have:
trace_idspan_idparent_span_idnamestart_timeend_timestatusattributes/ metadata
What to log in each step
1) Prompt build
Store:
- system/developer/user messages
- retrieved snippets
- prompt template version
- conversation history reference
- final assembled prompt
Be careful with secrets and PII. Often you store:
- full prompt in secure storage
- redacted prompt in normal logs
2) Model call
Store:
- provider: OpenAI / Anthropic / local vLLM / etc.
- model ID
- parameters
- request size
- response text
- tool call chunks / function-call outputs
- usage stats
- latency
- retry count
3) Tool use
Store:
- tool name
- args
- returned data
- whether tool succeeded
- duration
- any exception
- correlation to the model step that requested it
4) Final response
Store:
- final answer
- sources/citations
- structured output JSON
- whether it was streamed
- formatting/parsing errors
Common implementation approaches
Option A: Manual structured logging
You add logs yourself around each operation.
Example fields:
{
"trace_id": "abc123",
"span": "llm_call",
"model": "gpt-4.1",
"prompt": "...",
"response": "...",
"latency_ms": 842
}
Pros:
- simple
- flexible
Cons:
- you must maintain it everywhere
Option B: OpenTelemetry
Use OpenTelemetry to create spans and export them to:
- Jaeger
- Grafana Tempo
- Datadog
- Honeycomb
- New Relic
- OpenSearch, etc.
This is ideal if you want standard tracing plus metrics/logs.
Option C: LLM observability tools
Tools built specifically for LLM tracing:
- LangSmith
- Langfuse
- Arize Phoenix
- Helicone
- Weights & Biases Weave
- TruLens
These often capture:
- prompt chains
- tool calls
- token usage
- datasets/evals
- latency and cost
Recommended data model
A practical event schema:
{
"trace_id": "t1",
"span_id": "s2",
"parent_span_id": "s1",
"event_type": "llm_call",
"name": "generate_answer",
"timestamp": "2026-10-09T12:00:00Z",
"status": "ok",
"attributes": {
"model": "gpt-4.1",
"temperature": 0.2,
"input_tokens": 1200,
"output_tokens": 180,
"latency_ms": 910
}
}
Important guardrails
1) Redact sensitive data
Prompts often contain:
- personal data
- API keys
- secrets
- internal docs
Use:
- PII redaction
- secret detection
- allowlists for logged fields
- separate secure storage for raw content
2) Capture prompt/version metadata
If your prompt changes, you need:
- template version
- code version / git SHA
- model version
- retrieval index version
This makes debugging reproducible.
3) Trace retries and branching
LLM apps often have:
- retries
- fallback models
- conditional tool execution
- parallel tool calls
Your traces should preserve the execution graph, not just a flat log.
Example flow
User asks a question:
- Create
trace_id=xyz - Build prompt from system + user + retrieved docs
- Call model
- Model requests tool
search_docs - Run tool and record output
- Call model again with tool result
- Produce final answer
- Store all spans under same trace
If you want a quick start
If you’re building this yourself, start with:
- OpenTelemetry for spans
- JSON logging for payloads
- A tracing UI/backend like Jaeger or Langfuse
- Redaction middleware
Minimal checklist
Track:
- trace ID
- prompt inputs
- model name and params
- tool name and args
- outputs
- latency
- errors
- token usage
- prompt/model version
If you want, I can also give you:
- a Python example using OpenTelemetry,
- a LangChain/LangGraph tracing setup, or
- a production-ready event schema for LLM observability.