Prompt
How do I find the root cause of hallucinations in a RAG system?
Latest observation
To find the root cause of hallucinations in a RAG system, you want to debug the pipeline end-to-end and isolate which stage is failing:
1) First, classify the hallucination
Ask: Did the model make something up because it never saw the right info, or because it ignored the info it did see?
Common buckets:
- Retrieval failure: relevant documents were never retrieved.
- Context selection failure: retrieved docs were relevant, but the wrong chunks were included or ranked poorly.
- Generation failure: the model had the right evidence but still answered incorrectly, over-generalized, or hallucinated.
- Knowledge conflict: retrieved docs contradict each other or are stale.
- Prompt/format failure: the model wasn’t instructed strongly enough to ground answers in context.
2) Inspect the retrieval stage
Check whether the system is actually fetching evidence that supports the answer.
Questions to ask
- Was the query understood correctly?
- Are the top-k retrieved chunks relevant?
- Is the right source even in the index?
- Are embeddings and chunking appropriate for the domain?
Useful diagnostics
- Recall@k / hit rate: Does the correct source appear in top-k?
- Manual retrieval review: For hallucinated answers, inspect the retrieved chunks.
- Query rewrite logs: Did an LLM rewrite the query in a way that hurt retrieval?
- Metadata filters: Did filters exclude the correct document?
Common retrieval root causes
- Chunk size too large/small
- Poor embedding model
- Bad chunk boundaries
- Missing synonym/domain vocabulary
- Overly aggressive metadata filters
- Stale or incomplete index
3) Inspect the context window given to the LLM
Even if retrieval succeeds, the model may not “see” the evidence clearly.
Check:
- Are the best chunks actually included in the final prompt?
- Is the context too long and the key evidence buried?
- Are duplicates or irrelevant chunks crowding out useful ones?
- Is reranking working?
Common context issues
- Top-k too low
- Bad reranker
- Truncation of evidence due to context limits
- Multi-hop questions needing evidence from multiple chunks, but only one is included
- Lost provenance: evidence is present but not clearly attributed
4) Test generation behavior separately
If retrieval looks good, the issue may be the LLM itself.
Run these tests:
- Answer from context only: Force the model to answer strictly from provided text.
- Citation check: Require each claim to cite a source chunk.
- Abstention test: Ask it to say “I don’t know” if evidence is insufficient.
- Golden context test: Provide the exact supporting passage and see if it still hallucinates.
Generation failure patterns
- Model is too “helpful” and fills gaps
- Weak grounding prompt
- Temperature too high
- No penalty for unsupported claims
- Conflicts between system prompt and user instruction
5) Build a root-cause workflow
For each hallucinated response, log:
- User query
- Retrieved documents/chunks
- Reranker scores
- Final prompt sent to the model
- Final answer
- Citations or source IDs
- Whether the answer is supported by retrieved text
Then categorize:
- No relevant evidence retrieved
- Relevant evidence retrieved but excluded
- Relevant evidence included but ignored
- Evidence insufficient / ambiguous
- Contradictory evidence
- Wrong answer despite correct evidence
This makes debugging much faster.
6) Use evaluation sets
Create a small labeled set of questions with known answers and supporting sources.
Measure:
- Retrieval recall
- Answer correctness
- Faithfulness / groundedness
- Citation accuracy
- Abstention accuracy
This helps you see whether the issue is improving after changes.
7) Typical fixes by root cause
Retrieval problems
- Better chunking
- Better embeddings
- Hybrid search (BM25 + vector)
- Query rewriting
- Better reranking
- Index refresh / deduplication
Context problems
- Increase top-k
- Improve prompt packing
- Reduce irrelevant chunks
- Summarize long docs with provenance
- Ensure key passages survive truncation
Generation problems
- Lower temperature
- Stronger grounding prompt
- Force citations
- Constrain to answer only from context
- Add abstain behavior
Data problems
- Fix stale docs
- Remove conflicting sources
- Improve metadata quality
- Add source prioritization rules
8) A practical debugging trick
Take one hallucinated answer and run three controlled experiments:
- Retrieve-only: inspect top results without generation
- Oracle context: provide the exact correct passage to the model
- No-context baseline: see what the model says without retrieval
Interpretation:
- Fails in retrieve-only → retrieval issue
- Succeeds with oracle context but fails with real context → chunking/reranking/context issue
- Still hallucinates with oracle context → generation/prompt issue
If you want, I can also give you a step-by-step debugging checklist, or a template for logging and diagnosing hallucinations in production.