Prompt
I'm building a RAG chatbot, how do I inspect citations and source chunks?
Latest observation
To inspect citations and source chunks in a RAG chatbot, the key is to make your retrieval and generation pipeline observable end-to-end.
What to log/store
For each user question, capture:
- User query
- Retrieved chunks
- chunk text
- document/file name
- chunk ID
- page/section/URL
- retrieval score / rank
- Prompt sent to the LLM
- Model response
- Citations in the response
- which answer sentence maps to which chunk/source
- citation labels like
[1],[2]
- Metadata
- timestamp
- embedding model
- retriever config
- reranker config
- conversation/session ID
Best practice: assign stable source IDs
When you chunk documents, give every chunk a stable identifier, for example:
doc_042#chunk_003employee_handbook:p12:c2https://example.com/policy#section-4
Then every retrieved chunk can be referenced in the answer and in your logs.
Make citations explicit in the LLM prompt
Ask the model to cite from provided sources only. Example:
Use only the sources below. When you use a source, cite it as [1], [2], etc.
If the answer is not in the sources, say you don't know.
Sources:
[1] Chunk text...
[2] Chunk text...
This makes it easier to inspect why the model answered the way it did.
Return structured retrieval data
Instead of only returning plain text, return something like:
{
"answer": "The policy allows ... [1]",
"citations": [
{
"label": "[1]",
"chunk_id": "employee_handbook:p12:c2",
"document": "employee_handbook.pdf",
"page": 12,
"text": "..."
}
],
"retrieved_chunks": [
{
"chunk_id": "employee_handbook:p12:c2",
"score": 0.87,
"text": "..."
}
]
}
How to inspect source chunks during debugging
1. Print the retrieved chunks
For every query, inspect top-k results:
- Are they actually relevant?
- Are chunks too large or too small?
- Are duplicates being retrieved?
2. Compare citations to sources
Check whether each citation points to:
- the right chunk
- the right document
- the right page/section
3. Trace answer spans
If you want stronger grounding, map answer sentences to source spans:
- sentence-level attribution
- token-level highlighting
- quote extraction
4. Use a trace UI
Tools that help:
- LangSmith
- OpenTelemetry-based tracing
- Arize/Phoenix
- Weights & Biases tracing
- custom dashboard in your app
These let you click a response and inspect:
- query
- retriever output
- reranker output
- prompt
- final answer
Simple debugging workflow
- Ask a question.
- Look at the retrieved chunks.
- Verify the chunks contain the facts needed.
- Check whether the prompt includes those chunks.
- Inspect the citations in the answer.
- If the answer is wrong, determine whether the issue is:
- retrieval failure
- chunking issue
- prompt issue
- generation hallucination
- citation formatting issue
If you want higher-quality citations
Use:
- chunk-level citations
- reranking
- source quotation
- answer grounding checks
- “answer only from context” prompting
- post-generation citation validation
Recommended minimum setup
At minimum, your app should save:
querytop_k retrieved chunkssource metadatafinal answercitation mapping
If you want, I can show you:
- a LangChain example,
- a LlamaIndex example, or
- a custom JSON logging format for citations and source chunks.