Prompt

How do I trace a RAG request end to end from user input to cited answer?

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

To trace a RAG request end to end, you want a single request ID / trace ID that follows the query through every stage:

  1. User input received
  2. Query preprocessing
  3. Retrieval
  4. Reranking / filtering
  5. Prompt assembly
  6. LLM generation
  7. Citation mapping
  8. Final response returned

Here’s a practical way to do it.


1) Assign a trace ID at the entry point

When the user submits a question, create or propagate a trace_id and request_id.

Example:

  • trace_id: for the whole distributed trace
  • request_id: for the app-level request

Log the raw user query with these IDs.

{
  "trace_id": "abc123",
  "request_id": "req789",
  "event": "user_query_received",
  "query": "What are the refund terms?"
}

2) Log each RAG stage as a span

Treat the RAG pipeline like a distributed trace:

  • query_parse
  • rewrite_query
  • embed_query
  • retrieve_candidates
  • rerank_results
  • build_prompt
  • llm_generate
  • postprocess_answer
  • attach_citations

Each span should record:

  • start/end time
  • inputs/outputs
  • metadata
  • top-k retrieved docs and scores
  • prompt version
  • model name
  • token counts

Example span log:

{
  "trace_id": "abc123",
  "span": "retrieve_candidates",
  "top_k": 5,
  "results": [
    {"doc_id": "doc1", "score": 0.92},
    {"doc_id": "doc8", "score": 0.87}
  ]
}

3) Preserve document identity through the pipeline

Every chunk you retrieve should have stable metadata:

  • doc_id
  • chunk_id
  • source_uri
  • title
  • page
  • offsets
  • version

This is crucial because citations depend on being able to map output text back to source chunks.

Example chunk metadata:

{
  "doc_id": "policy_2025",
  "chunk_id": "policy_2025#p4c2",
  "source_uri": "s3://kb/policy.pdf",
  "page": 4,
  "score": 0.91
}

4) Log the exact prompt sent to the model

To understand why the model answered a certain way, store:

  • system prompt
  • developer prompt
  • retrieved context chunks
  • formatting instructions
  • query

Best practice:

  • save the prompt template version
  • save the exact rendered prompt
  • redact sensitive data if needed

Example:

{
  "trace_id": "abc123",
  "event": "prompt_built",
  "prompt_version": "rag-v12",
  "context_chunks": ["policy_2025#p4c2", "policy_2025#p6c1"]
}

5) Capture retrieval scores and selection logic

If you do any filtering or reranking, log:

  • initial candidates
  • reranker scores
  • thresholds
  • which docs were dropped and why

This helps explain citation gaps or missing answers.

Example:

{
  "trace_id": "abc123",
  "event": "rerank_complete",
  "kept": ["doc1#chunk3", "doc8#chunk1"],
  "dropped": [
    {"chunk_id": "doc2#chunk4", "reason": "score_below_threshold"}
  ]
}

6) Track citation provenance explicitly

If you want cited answers, you need a mapping from generated answer spans to source chunks.

A strong pattern is:

  • retrieve chunks
  • ask the model to answer with inline citation markers like [1], [2]
  • map those markers back to chunk IDs in your app

Example response format:

  • [1] → policy_2025#p4c2
  • [2] → policy_2025#p6c1

Then log:

{
  "trace_id": "abc123",
  "event": "citation_map",
  "citations": {
    "1": "policy_2025#p4c2",
    "2": "policy_2025#p6c1"
  }
}

7) Store final answer with provenance

Your final log should include:

  • final answer
  • citations
  • source chunks
  • model output metadata
  • latency by stage

Example:

{
  "trace_id": "abc123",
  "event": "final_answer",
  "answer": "Refunds are allowed within 30 days [1]. Exceptions apply for digital goods [2].",
  "citations": [
    {"marker": "1", "chunk_id": "policy_2025#p4c2"},
    {"marker": "2", "chunk_id": "policy_2025#p6c1"}
  ]
}

8) Use observability tooling

For real end-to-end tracing, use a tracing system:

  • OpenTelemetry
  • Jaeger
  • Tempo
  • Datadog APM
  • New Relic

In RAG, each pipeline step can be a span under one trace.

Typical trace tree:

  • HTTP POST /chat
    • parse_request
    • embed_query
    • vector_search
    • rerank
    • assemble_prompt
    • llm_call
    • citation_postprocess
    • response_send

9) Add debugging views

A useful debugging UI should show:

  • user question
  • rewritten query
  • retrieved documents/chunks
  • scores
  • final prompt
  • raw model output
  • final answer
  • citation mapping
  • timing breakdown

This makes it much easier to diagnose:

  • bad retrieval
  • hallucinations
  • wrong citations
  • prompt injection
  • latency bottlenecks

10) Recommended minimum data to persist

At minimum, persist these fields per request:

  • trace_id
  • user_query
  • rewritten_query if any
  • retrieved chunk IDs and scores
  • prompt template version
  • exact context included
  • model name and parameters
  • raw model output
  • final answer
  • citation map
  • stage timings

Example end-to-end trace

{
  "trace_id": "abc123",
  "request": {
    "query": "What are the refund terms?"
  },
  "stages": [
    {
      "name": "retrieve_candidates",
      "results": [
        {"chunk_id": "policy_2025#p4c2", "score": 0.92},
        {"chunk_id": "policy_2025#p6c1", "score": 0.87}
      ]
    },
    {
      "name": "build_prompt",
      "prompt_version": "rag-v12"
    },
    {
      "name": "llm_generate",
      "model": "gpt-4.1"
    }
  ],
  "answer": {
    "text": "Refunds are allowed within 30 days [1]. Digital goods are excluded [2].",
    "citations": {
      "1": "policy_2025#p4c2",
      "2": "policy_2025#p6c1"
    }
  }
}

Best practice summary

If you only do five things, do these:

  1. Use one trace ID across the full request
  2. Log every stage as a span
  3. Keep stable chunk/document IDs
  4. Store the exact prompt and retrieved context
  5. Record a citation map from answer markers to source chunks

If you want, I can also give you:

  • a sample OpenTelemetry tracing schema for RAG
  • a Python implementation
  • or a LangChain/LlamaIndex tracing example

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.