Prompt

How do I trace a user complaint through an LLM app from prompt to tool calls to final response?

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

A good way to trace a complaint through an LLM app is to treat one user request as a single trace and record every step in that trace:

  1. Incoming user message
  2. Prompt construction
  3. Model call(s)
  4. Tool/function calls
  5. Tool outputs
  6. Final model response
  7. Any retries, errors, or fallback paths

Here’s a practical approach.


1) Assign a trace ID to each user request

When a user sends a message, create a unique trace_id and attach it to everything downstream.

Example metadata:

  • trace_id
  • user_id
  • session_id
  • request_id
  • message_id

This lets you connect:

  • the original complaint
  • the prompt sent to the LLM
  • any tool calls
  • the final answer

2) Log the prompt in structured form

Don’t just log “the prompt.” Log the components separately:

  • system prompt
  • developer prompt
  • conversation history
  • retrieved context
  • user message
  • tool instructions / schemas

Example structure:

{
  "trace_id": "abc123",
  "stage": "prompt_build",
  "system_prompt": "...",
  "developer_prompt": "...",
  "conversation": [
    {"role": "user", "content": "My order is missing items"}
  ],
  "retrieved_context": [
    {"source": "order_db", "content": "Order #456 shipped yesterday"}
  ]
}

This helps answer:

  • What did the model actually see?
  • Was the complaint already misrepresented before the model ran?

3) Capture the exact model input and output

For every LLM call, record:

  • model name/version
  • parameters (temperature, top_p, etc.)
  • full input messages
  • raw model output
  • token usage
  • latency
  • any stop reasons

If your model supports tool calling, store the assistant message exactly as returned.

Example:

{
  "trace_id": "abc123",
  "stage": "llm_call",
  "model": "gpt-4.1",
  "temperature": 0.2,
  "input_messages": [...],
  "output": {
    "role": "assistant",
    "content": null,
    "tool_calls": [
      {
        "name": "lookup_order",
        "arguments": "{\"order_id\":\"456\"}"
      }
    ]
  }
}

4) Log each tool call as a separate span

Treat every tool call as its own event/span in the trace.

Record:

  • tool name
  • arguments
  • start/end time
  • success/failure
  • output payload
  • error details

Example:

{
  "trace_id": "abc123",
  "stage": "tool_call",
  "tool": "lookup_order",
  "arguments": {"order_id": "456"},
  "result": {
    "status": "ok",
    "data": {
      "shipment_status": "delivered",
      "missing_items": ["charger"]
    }
  }
}

This is critical because many user complaints are caused by:

  • bad tool arguments
  • wrong tool chosen
  • stale tool data
  • tool failures that get papered over by the model

5) Trace the full chain, not just the final answer

A complaint often comes from a mismatch anywhere in the chain:

  • user said “refund never arrived”
  • prompt omitted the refund tool
  • model inferred wrong order ID
  • tool returned partial data
  • final answer incorrectly claimed refund was processed

So keep a parent-child structure like this:

  • trace_id
    • span: prompt_build
    • span: llm_call_1
      • span: tool_call_lookup_order
      • span: tool_call_check_refund
    • span: llm_call_2
    • span: final_response

This lets you reconstruct the path end to end.


6) Add correlation IDs to your application logs

Make sure your app logs, tool service logs, and database logs all include the same identifiers.

Recommended fields:

  • trace_id
  • span_id
  • parent_span_id
  • request_id
  • conversation_id
  • user_id

If you use distributed tracing tools, propagate trace context across:

  • API gateway
  • app server
  • retrieval service
  • tool backend
  • LLM provider wrapper

7) Store rendered prompts and raw prompts separately

Useful distinction:

  • Rendered prompt: exact text/messages sent to the model
  • Template inputs: variables used to fill the prompt
  • Template version: which prompt template was used

This helps catch issues like:

  • prompt template regression
  • bad variable interpolation
  • accidental truncation

8) Capture retries and branch decisions

If your app retries a model call or takes different paths based on output, log that too.

Examples:

  • model response invalid JSON → retry
  • tool returned error → fallback answer
  • confidence low → ask user a clarifying question

Record:

  • why the branch happened
  • what input caused it
  • which output was chosen

9) Use a trace viewer or observability stack

Common approach:

  • OpenTelemetry for traces
  • logs in JSON
  • traces in Jaeger, Grafana Tempo, Datadog, Honeycomb, etc.

You want to be able to click one trace_id and see:

  • the complaint
  • the prompt
  • every model invocation
  • each tool call
  • final response

10) Build a “complaint replay” workflow

For debugging user complaints, create a replay tool that can show:

  1. original user message
  2. full prompt context
  3. model inputs/outputs
  4. tool inputs/outputs
  5. final answer
  6. diffs between expected and actual behavior

This is especially useful for:

  • support investigations
  • QA
  • regression testing
  • incident reviews

11) Be careful with privacy and secrets

Don’t blindly log everything. Redact:

  • passwords
  • API keys
  • payment data
  • personal data if not needed

Good practice:

  • structured logs with redaction rules
  • encrypted storage
  • access controls
  • short retention where appropriate

12) Minimal example tracing flow

A practical trace might look like:

trace_id=abc123
  user_input: "I was charged twice"

  prompt_build:
    system_prompt_version=v12
    retrieved_context=[billing policy...]

  llm_call_1:
    output=call tool "find_transactions" with customer_id=789

  tool_call_find_transactions:
    result=[txn1, txn2]

  llm_call_2:
    output=final response "I found two charges..."

  final_response:
    "I found two charges and one appears duplicate..."

That gives you a clear path from complaint to response.


Recommended implementation pattern

If you’re building this from scratch, use:

  • OpenTelemetry spans for each stage
  • structured JSON logs for prompt/tool payloads
  • trace context propagation across services
  • prompt versioning
  • tool call auditing
  • redaction

If you want a very practical debugging checklist

When a user complains, inspect in order:

  1. Was the original user message captured correctly?
  2. Which prompt template/version was used?
  3. What context was injected?
  4. What did the model actually receive?
  5. Did it call the right tools?
  6. Were the tool arguments correct?
  7. What did the tools return?
  8. Did the model summarize tool output accurately?
  9. Was there a retry/fallback that changed behavior?
  10. Where exactly did the output diverge from expected?

If you want, I can also give you:

  • a sample OpenTelemetry schema for LLM traces
  • a Python middleware example
  • or a JSON log format you can drop into your app.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.