Prompt
How do I measure hallucinations in a chatbot response?
Latest observation
To measure hallucinations in a chatbot response, you want to check whether the answer is grounded in trusted sources and whether any claims are unsupported, incorrect, or fabricated.
1) Define what “hallucination” means for your use case
A hallucination is usually any of these:
- Unsupported claim: the chatbot states something not found in the source material
- Incorrect claim: the statement conflicts with the source or known facts
- Fabricated detail: a made-up name, date, citation, quote, or event
- Overconfident uncertainty: the chatbot presents uncertain info as certain
If your chatbot is using retrieval (RAG), the key question is:
“Is each answer statement supported by the retrieved context?”
2) Use claim-level evaluation
Break the response into individual claims, then check each one against a reference source.
Example response:
“The product launched in 2022, supports SSO, and is available in Europe.”
Claims:
- Product launched in 2022
- Supports SSO
- Available in Europe
Label each claim as:
- Supported
- Contradicted
- Not enough evidence
Then compute:
- Hallucination rate = contradicted or unsupported claims / total claims
- Faithfulness rate = supported claims / total claims
3) Common metrics
For fact-grounded systems
- Precision of factual claims: how many statements are correct
- Faithfulness / groundedness: how much of the answer is supported by the provided context
- Hallucination rate: fraction of unsupported or false claims
- Citation accuracy: whether cited sources actually support the claim
For retrieval-based systems
- Context precision: how much retrieved context is relevant
- Context recall: whether the retrieved context contains the needed facts
- Answer faithfulness: whether the answer uses only retrieved facts
4) How to measure it in practice
A. Human evaluation
Have annotators rate each response:
- Supported
- Unsupported
- Contradicted
- Unclear
This is the most reliable method, especially for nuanced answers.
B. Automated evaluation
Use:
- LLM-as-judge: another model checks if each claim is supported by context
- NLI/entailment models: determine whether context entails the answer
- Fact-checking pipelines: extract claims and verify them against sources
Automated methods are cheaper, but less reliable than humans.
C. Reference-based testing
If you have a gold answer:
- Compare the chatbot response to the reference
- Measure semantic similarity, factual overlap, or exact correctness
This works well for structured tasks, but less so for open-ended chat.
5) A simple scoring scheme
You can score each response like this:
- 0 = fully grounded, no unsupported claims
- 1 = minor hallucination, mostly correct
- 2 = moderate hallucination, several unsupported claims
- 3 = severe hallucination, largely fabricated
Or use percentages:
- 100% supported
- 75% supported, 25% unsupported
- etc.
6) Important caveats
- A response can be useful but partially hallucinated
- Some claims may be true in the real world but not supported by your provided source
- Ambiguous questions make evaluation harder
- “Hallucination” depends on the required standard: general knowledge vs. source-grounded answers
7) Practical recommendation
If you want a straightforward method:
- Split response into atomic claims
- Compare each claim to trusted source material
- Label claims as supported / unsupported / contradicted
- Compute hallucination rate
- Track results over many conversations
If you want, I can also give you:
- a rubric for human annotators
- a Python script to score hallucinations
- or a prompt for an LLM judge to evaluate responses automatically.