Prompt
How do I trace a user complaint through an LLM app from prompt to tool calls to final response?
Latest observation
A good way to trace a complaint through an LLM app is to treat one user request as a single trace and record every step in that trace:
- Incoming user message
- Prompt construction
- Model call(s)
- Tool/function calls
- Tool outputs
- Final model response
- Any retries, errors, or fallback paths
Here’s a practical approach.
1) Assign a trace ID to each user request
When a user sends a message, create a unique trace_id and attach it to everything downstream.
Example metadata:
trace_iduser_idsession_idrequest_idmessage_id
This lets you connect:
- the original complaint
- the prompt sent to the LLM
- any tool calls
- the final answer
2) Log the prompt in structured form
Don’t just log “the prompt.” Log the components separately:
- system prompt
- developer prompt
- conversation history
- retrieved context
- user message
- tool instructions / schemas
Example structure:
{
"trace_id": "abc123",
"stage": "prompt_build",
"system_prompt": "...",
"developer_prompt": "...",
"conversation": [
{"role": "user", "content": "My order is missing items"}
],
"retrieved_context": [
{"source": "order_db", "content": "Order #456 shipped yesterday"}
]
}
This helps answer:
- What did the model actually see?
- Was the complaint already misrepresented before the model ran?
3) Capture the exact model input and output
For every LLM call, record:
- model name/version
- parameters (
temperature,top_p, etc.) - full input messages
- raw model output
- token usage
- latency
- any stop reasons
If your model supports tool calling, store the assistant message exactly as returned.
Example:
{
"trace_id": "abc123",
"stage": "llm_call",
"model": "gpt-4.1",
"temperature": 0.2,
"input_messages": [...],
"output": {
"role": "assistant",
"content": null,
"tool_calls": [
{
"name": "lookup_order",
"arguments": "{\"order_id\":\"456\"}"
}
]
}
}
4) Log each tool call as a separate span
Treat every tool call as its own event/span in the trace.
Record:
- tool name
- arguments
- start/end time
- success/failure
- output payload
- error details
Example:
{
"trace_id": "abc123",
"stage": "tool_call",
"tool": "lookup_order",
"arguments": {"order_id": "456"},
"result": {
"status": "ok",
"data": {
"shipment_status": "delivered",
"missing_items": ["charger"]
}
}
}
This is critical because many user complaints are caused by:
- bad tool arguments
- wrong tool chosen
- stale tool data
- tool failures that get papered over by the model
5) Trace the full chain, not just the final answer
A complaint often comes from a mismatch anywhere in the chain:
- user said “refund never arrived”
- prompt omitted the refund tool
- model inferred wrong order ID
- tool returned partial data
- final answer incorrectly claimed refund was processed
So keep a parent-child structure like this:
trace_idspan: prompt_buildspan: llm_call_1span: tool_call_lookup_orderspan: tool_call_check_refund
span: llm_call_2span: final_response
This lets you reconstruct the path end to end.
6) Add correlation IDs to your application logs
Make sure your app logs, tool service logs, and database logs all include the same identifiers.
Recommended fields:
trace_idspan_idparent_span_idrequest_idconversation_iduser_id
If you use distributed tracing tools, propagate trace context across:
- API gateway
- app server
- retrieval service
- tool backend
- LLM provider wrapper
7) Store rendered prompts and raw prompts separately
Useful distinction:
- Rendered prompt: exact text/messages sent to the model
- Template inputs: variables used to fill the prompt
- Template version: which prompt template was used
This helps catch issues like:
- prompt template regression
- bad variable interpolation
- accidental truncation
8) Capture retries and branch decisions
If your app retries a model call or takes different paths based on output, log that too.
Examples:
- model response invalid JSON → retry
- tool returned error → fallback answer
- confidence low → ask user a clarifying question
Record:
- why the branch happened
- what input caused it
- which output was chosen
9) Use a trace viewer or observability stack
Common approach:
- OpenTelemetry for traces
- logs in JSON
- traces in Jaeger, Grafana Tempo, Datadog, Honeycomb, etc.
You want to be able to click one trace_id and see:
- the complaint
- the prompt
- every model invocation
- each tool call
- final response
10) Build a “complaint replay” workflow
For debugging user complaints, create a replay tool that can show:
- original user message
- full prompt context
- model inputs/outputs
- tool inputs/outputs
- final answer
- diffs between expected and actual behavior
This is especially useful for:
- support investigations
- QA
- regression testing
- incident reviews
11) Be careful with privacy and secrets
Don’t blindly log everything. Redact:
- passwords
- API keys
- payment data
- personal data if not needed
Good practice:
- structured logs with redaction rules
- encrypted storage
- access controls
- short retention where appropriate
12) Minimal example tracing flow
A practical trace might look like:
trace_id=abc123
user_input: "I was charged twice"
prompt_build:
system_prompt_version=v12
retrieved_context=[billing policy...]
llm_call_1:
output=call tool "find_transactions" with customer_id=789
tool_call_find_transactions:
result=[txn1, txn2]
llm_call_2:
output=final response "I found two charges..."
final_response:
"I found two charges and one appears duplicate..."
That gives you a clear path from complaint to response.
Recommended implementation pattern
If you’re building this from scratch, use:
- OpenTelemetry spans for each stage
- structured JSON logs for prompt/tool payloads
- trace context propagation across services
- prompt versioning
- tool call auditing
- redaction
If you want a very practical debugging checklist
When a user complains, inspect in order:
- Was the original user message captured correctly?
- Which prompt template/version was used?
- What context was injected?
- What did the model actually receive?
- Did it call the right tools?
- Were the tool arguments correct?
- What did the tools return?
- Did the model summarize tool output accurately?
- Was there a retry/fallback that changed behavior?
- Where exactly did the output diverge from expected?
If you want, I can also give you:
- a sample OpenTelemetry schema for LLM traces
- a Python middleware example
- or a JSON log format you can drop into your app.