Prompt
What should I use to score hallucinations in LLM answers?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
A good way to score hallucinations in LLM answers is to use evidence-based factuality metrics rather than generic similarity metrics.
Best options by use case
1. If you have source documents / retrieved context
Use faithfulness / groundedness scoring:
- Check whether each claim in the answer is supported by the provided context.
- Common choices:
- RAGAS Faithfulness
- TruLens Groundedness
- Faithfulness / context precision-recall style metrics
- LLM-as-judge with citations/evidence checking
This is usually the best approach for RAG systems.
2. If you want to measure factual hallucination against a reference answer
Use factual accuracy metrics:
- FactScore: breaks answers into atomic facts and verifies them
- QAFactEval: question-answer based factual consistency
- SummaC: often used for summarization factual consistency
- BERTScore / ROUGE are not good hallucination metrics by themselves, because they measure overlap, not truth.
3. If you need a practical production metric
Use a combination of:
- Claim extraction
- Entailment / verification against evidence
- LLM judge with structured rubric
- Human spot checks for calibration
A common scoring setup is:
- Split answer into atomic claims
- Retrieve supporting evidence
- Score each claim as:
- supported
- contradicted
- not enough evidence
- Hallucination score = fraction unsupported/contradicted claims
Recommended default
If you want one answer:
- For RAG / grounded answers: use RAGAS Faithfulness
- For general factual hallucination: use FactScore
- For quick evaluation: use an LLM-as-judge groundedness rubric, but validate it against human labels
What to avoid
Avoid using these as hallucination scores:
- ROUGE
- BLEU
- exact match
- embedding similarity alone
They can reward fluent but false answers.
Practical metric definition
A simple hallucination score can be:
[ \text{Hallucination Rate} = \frac{\text{unsupported claims}}{\text{total claims}} ]
or
[ \text{Hallucination Score} = 1 - \text{Faithfulness} ]
If you want, I can recommend a specific metric stack for:
- RAG systems
- summarization
- chatbots
- medical/legal QA