Prompt
I'm building a RAG assistant and want to track citations, hallucinations, and retrieval quality
Latest observation
A good way to think about this is to instrument your RAG system at three layers:
- Retrieval quality — did you fetch the right evidence?
- Grounding / hallucination — is the answer supported by that evidence?
- Citations — can you attribute each claim to a source chunk?
Below is a practical framework you can implement.
1) What to log for every RAG request
Store a single “trace” record per user query with:
- query
- retrieved_chunks: ids, text, source metadata, rank, scores
- prompt_sent_to_llm
- model_output
- citations_emitted: claim → chunk ids
- final_answer
- latency breakdown
- user feedback if available
A useful unit is a claim-level trace, not just a whole-answer trace.
Suggested trace schema
{
"request_id": "uuid",
"timestamp": "2026-10-09T12:34:56Z",
"query": "What is the refund policy?",
"retrieval": {
"top_k": 5,
"chunks": [
{
"chunk_id": "doc12#p3",
"source": "help_center/refunds.md",
"rank": 1,
"score": 0.87,
"text": "Refunds are available within 30 days..."
}
]
},
"generation": {
"answer": "You can request a refund within 30 days ...",
"claims": [
{
"claim_id": "c1",
"text": "Refunds are available within 30 days.",
"citations": ["doc12#p3"]
}
]
},
"metrics": {
"retrieval_recall_at_k": 0.8,
"faithfulness_score": 0.92,
"citation_precision": 1.0
}
}
2) Citation tracking: make the model cite at claim level
To avoid “one citation for the whole paragraph,” require the model to emit:
- Answer text
- Structured claims
- Per-claim citations
- Optional: quote spans from retrieved chunks
Example output format
{
"answer": "Refunds are available within 30 days of purchase.",
"claims": [
{
"text": "Refunds are available within 30 days of purchase.",
"citations": [
{
"chunk_id": "doc12#p3",
"quote": "Refunds are available within 30 days of purchase."
}
]
}
]
}
Best practices
- Force citations on every factual sentence.
- Prefer chunk ids over URLs alone.
- Keep chunk text and offsets so you can verify exact support.
- If unsupported, allow the model to say: “I couldn’t find evidence for that.”
3) Hallucination tracking: measure whether claims are supported
Hallucination in RAG usually means the answer contains claims that are not grounded in retrieved evidence.
Practical metrics
Track at the claim level:
- Supported claim rate = supported claims / total claims
- Unsupported claim rate = unsupported claims / total claims
- Attribution accuracy = cited chunk actually supports the claim
- Abstention rate = how often the model correctly says it doesn’t know
How to detect support
You can use a few approaches:
A. Rule-based / lexical overlap
Good for quick checks:
- Extract named entities, numbers, dates, policy terms
- Compare against retrieved chunk text
Pros: cheap, fast
Cons: weak for paraphrases
B. LLM-based entailment judge
Prompt a separate model:
- “Does this evidence entail the claim?”
- Output: supported / contradicted / not enough info
Pros: much better coverage
Cons: more cost, needs calibration
C. Embedding similarity
Useful as a signal, not a final decision:
- claim embedding vs retrieved chunk embeddings
Pros: simple
Cons: not enough to prove support
Recommended hybrid
Use:
- lexical checks for numbers/dates/entities
- entailment judge for final support decision
- similarity only as a retrieval diagnostic
4) Retrieval quality: measure whether the right evidence was found
Retrieval quality should be evaluated independently of generation.
Core metrics
At query level:
- Recall@k: did any relevant chunk appear in top k?
- MRR: how early did the first relevant chunk appear?
- nDCG@k: ranking quality with graded relevance
- Precision@k: how many retrieved chunks were actually relevant?
- Coverage: are all needed subtopics represented?
If you have labeled relevant chunks
Great — compute classic IR metrics.
If you don’t have labels
Create an evaluation set with:
- query
- gold supporting chunk(s)
- expected answer
- maybe gold facts
Then use:
- human labeling
- LLM-assisted labeling with human review
RAG-specific retrieval diagnostics
Also track:
- context utilization: did the model use the retrieved chunks?
- evidence leakage: did it answer from prior knowledge instead of context?
- chunk duplication: are top-k chunks redundant?
- source diversity: are you retrieving from multiple docs when needed?
5) Suggested end-to-end evaluation pipeline
Offline eval
For a benchmark set:
- Retrieve top-k
- Generate answer
- Extract claims
- Judge each claim as supported / unsupported / contradicted
- Score citations and retrieval
Metrics to report
- Retrieval: Recall@5, MRR, nDCG@10
- Generation: faithfulness, correctness, completeness
- Citation quality: citation precision, citation recall
- Hallucination: unsupported-claim rate
- Abstention: correct “I don’t know” rate
6) A simple scoring model you can implement
For each claim:
- Supported = 1 if at least one retrieved chunk entails it
- Cited correctly = 1 if the cited chunk supports it
- Hallucinated = 1 if unsupported but asserted as fact
Then compute:
citation_precision = supported_and_cited / cited_claims
citation_recall = supported_and_cited / supported_claims
hallucination_rate = unsupported_claims / total_claims
You can also compute a per-answer score:
faithfulness = supported_claims / total_claims
7) Implementation tips
Prompting
Tell the model:
- use only provided context
- cite every factual statement
- if evidence is insufficient, say so
- avoid broad unsupported claims
Chunking
Better citations start with better chunks:
- keep chunks semantically coherent
- include headings and metadata
- avoid overly large chunks
- preserve document structure
Logging
Log raw retrieved text and the exact prompt, because debugging citation issues without them is painful.
Human review
Create a sampling dashboard for:
- unsupported claims
- low-confidence retrieval
- answers with missing citations
- disagreements between model and judge
8) Tooling ideas
You can implement this with:
- OpenTelemetry for traces
- LangSmith / Arize Phoenix / W&B / Weights & Biases for RAG observability
- custom evaluation scripts for retrieval and faithfulness
- LLM-as-judge for claim support
9) Recommended minimal setup
If you want the smallest useful version:
- Log query, retrieved chunks, answer, citations
- Extract claims from the answer
- Use an entailment judge to label each claim
- Compute:
- retrieval Recall@k
- supported-claim rate
- citation precision
- hallucination rate
That gives you a solid baseline.
If you want, I can also provide:
- a Postgres schema for RAG traces,
- a Python evaluation script,
- or a LangChain/LlamaIndex instrumentation example.