Prompt
Do I need to evaluate both hallucinations and faithfulness?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Usually, yes — if you want a complete evaluation of a generated answer, you should look at both:
- Hallucinations: whether the model states things that are unsupported or false.
- Faithfulness: whether the answer stays grounded in the provided source/context.
Why both?
They’re related, but not identical:
- A response can be faithful to the source but still contain a subtle hallucination if it adds unsupported extra detail.
- A response can have no obvious hallucinations but still be unfaithful if it misrepresents the source, omits crucial caveats, or draws an invalid conclusion from the context.
Practical rule
- If your task is answering from provided documents/context, prioritize faithfulness.
- If your task is general factual generation, prioritize hallucination detection.
- For most RAG / grounded QA evaluations, it’s best to measure both, because together they give a clearer picture of quality.
If you want, I can also explain the difference between faithfulness, factuality, and hallucination in one simple table.