Prompt
How do I trace a RAG request end to end from user input to cited answer?
Latest observation
To trace a RAG request end to end, you want a single request ID / trace ID that follows the query through every stage:
- User input received
- Query preprocessing
- Retrieval
- Reranking / filtering
- Prompt assembly
- LLM generation
- Citation mapping
- Final response returned
Here’s a practical way to do it.
1) Assign a trace ID at the entry point
When the user submits a question, create or propagate a trace_id and request_id.
Example:
trace_id: for the whole distributed tracerequest_id: for the app-level request
Log the raw user query with these IDs.
{
"trace_id": "abc123",
"request_id": "req789",
"event": "user_query_received",
"query": "What are the refund terms?"
}
2) Log each RAG stage as a span
Treat the RAG pipeline like a distributed trace:
query_parserewrite_queryembed_queryretrieve_candidatesrerank_resultsbuild_promptllm_generatepostprocess_answerattach_citations
Each span should record:
- start/end time
- inputs/outputs
- metadata
- top-k retrieved docs and scores
- prompt version
- model name
- token counts
Example span log:
{
"trace_id": "abc123",
"span": "retrieve_candidates",
"top_k": 5,
"results": [
{"doc_id": "doc1", "score": 0.92},
{"doc_id": "doc8", "score": 0.87}
]
}
3) Preserve document identity through the pipeline
Every chunk you retrieve should have stable metadata:
doc_idchunk_idsource_urititlepageoffsetsversion
This is crucial because citations depend on being able to map output text back to source chunks.
Example chunk metadata:
{
"doc_id": "policy_2025",
"chunk_id": "policy_2025#p4c2",
"source_uri": "s3://kb/policy.pdf",
"page": 4,
"score": 0.91
}
4) Log the exact prompt sent to the model
To understand why the model answered a certain way, store:
- system prompt
- developer prompt
- retrieved context chunks
- formatting instructions
- query
Best practice:
- save the prompt template version
- save the exact rendered prompt
- redact sensitive data if needed
Example:
{
"trace_id": "abc123",
"event": "prompt_built",
"prompt_version": "rag-v12",
"context_chunks": ["policy_2025#p4c2", "policy_2025#p6c1"]
}
5) Capture retrieval scores and selection logic
If you do any filtering or reranking, log:
- initial candidates
- reranker scores
- thresholds
- which docs were dropped and why
This helps explain citation gaps or missing answers.
Example:
{
"trace_id": "abc123",
"event": "rerank_complete",
"kept": ["doc1#chunk3", "doc8#chunk1"],
"dropped": [
{"chunk_id": "doc2#chunk4", "reason": "score_below_threshold"}
]
}
6) Track citation provenance explicitly
If you want cited answers, you need a mapping from generated answer spans to source chunks.
A strong pattern is:
- retrieve chunks
- ask the model to answer with inline citation markers like
[1],[2] - map those markers back to chunk IDs in your app
Example response format:
[1]→policy_2025#p4c2[2]→policy_2025#p6c1
Then log:
{
"trace_id": "abc123",
"event": "citation_map",
"citations": {
"1": "policy_2025#p4c2",
"2": "policy_2025#p6c1"
}
}
7) Store final answer with provenance
Your final log should include:
- final answer
- citations
- source chunks
- model output metadata
- latency by stage
Example:
{
"trace_id": "abc123",
"event": "final_answer",
"answer": "Refunds are allowed within 30 days [1]. Exceptions apply for digital goods [2].",
"citations": [
{"marker": "1", "chunk_id": "policy_2025#p4c2"},
{"marker": "2", "chunk_id": "policy_2025#p6c1"}
]
}
8) Use observability tooling
For real end-to-end tracing, use a tracing system:
- OpenTelemetry
- Jaeger
- Tempo
- Datadog APM
- New Relic
In RAG, each pipeline step can be a span under one trace.
Typical trace tree:
HTTP POST /chatparse_requestembed_queryvector_searchrerankassemble_promptllm_callcitation_postprocessresponse_send
9) Add debugging views
A useful debugging UI should show:
- user question
- rewritten query
- retrieved documents/chunks
- scores
- final prompt
- raw model output
- final answer
- citation mapping
- timing breakdown
This makes it much easier to diagnose:
- bad retrieval
- hallucinations
- wrong citations
- prompt injection
- latency bottlenecks
10) Recommended minimum data to persist
At minimum, persist these fields per request:
trace_iduser_queryrewritten_queryif any- retrieved chunk IDs and scores
- prompt template version
- exact context included
- model name and parameters
- raw model output
- final answer
- citation map
- stage timings
Example end-to-end trace
{
"trace_id": "abc123",
"request": {
"query": "What are the refund terms?"
},
"stages": [
{
"name": "retrieve_candidates",
"results": [
{"chunk_id": "policy_2025#p4c2", "score": 0.92},
{"chunk_id": "policy_2025#p6c1", "score": 0.87}
]
},
{
"name": "build_prompt",
"prompt_version": "rag-v12"
},
{
"name": "llm_generate",
"model": "gpt-4.1"
}
],
"answer": {
"text": "Refunds are allowed within 30 days [1]. Digital goods are excluded [2].",
"citations": {
"1": "policy_2025#p4c2",
"2": "policy_2025#p6c1"
}
}
}
Best practice summary
If you only do five things, do these:
- Use one trace ID across the full request
- Log every stage as a span
- Keep stable chunk/document IDs
- Store the exact prompt and retrieved context
- Record a citation map from answer markers to source chunks
If you want, I can also give you:
- a sample OpenTelemetry tracing schema for RAG
- a Python implementation
- or a LangChain/LlamaIndex tracing example