Prompt
What should I use for evaluating answer quality in RAG?
Latest observation
For RAG, “answer quality” is usually best evaluated with a bundle of metrics, not a single one. The right choice depends on whether you care most about answer correctness, groundedness in retrieved context, retrieval quality, or user satisfaction.
Recommended core metrics
1. Answer correctness / relevance
Measures whether the final answer actually answers the user’s question.
- Exact Match (EM): strict string match; useful for factoid QA, but brittle.
- F1 / token overlap: better for short extractive answers.
- LLM-as-judge / rubric-based scoring: best for open-ended answers.
- Judge: “Does the answer fully address the question?”
- Score on a scale, e.g. 1–5.
2. Faithfulness / groundedness
Measures whether the answer is supported by the retrieved passages.
- Faithfulness / groundedness score:
- “Is every key claim supported by the context?”
- Very important in RAG because a fluent but unsupported answer can look good while being wrong.
3. Context relevance
Measures whether retrieved chunks are actually useful for answering the question.
- Context precision: how much of the retrieved context is relevant
- Context recall: did retrieval include the needed evidence?
- Context relevance score via LLM judge
4. Retrieval quality
If you’re tuning retrieval separately, use IR metrics:
- Recall@k: did the gold evidence appear in top k?
- MRR: how early the first relevant chunk appears
- nDCG@k: ranking quality when multiple relevant chunks exist
5. End-to-end task success
Best for product evaluation.
- Task completion / user satisfaction
- Human ratings on:
- correctness
- completeness
- clarity
- citation quality
- refusal behavior when evidence is missing
Best practical setup
If you want a strong default evaluation stack for RAG, use:
- Answer correctness — LLM judge or human rubric
- Faithfulness / groundedness — whether claims are supported by retrieved context
- Context recall@k — whether the needed evidence was retrieved
- Citation accuracy — if you provide citations, verify they support the claims
- Latency/cost — important for production tradeoffs
If you need one metric only
Use LLM-as-judge grounded answer quality, with a rubric that scores:
- correctness
- completeness
- groundedness
- clarity
That’s usually the most useful single metric for open-ended RAG answers.
Suggested rubric
Example 1–5 rubric:
- 5: fully correct, complete, and fully supported by context
- 4: mostly correct, minor omissions, supported
- 3: partially correct or partially supported
- 2: major issues in correctness or grounding
- 1: incorrect or hallucinated
Tooling
Common choices:
- RAGAS: popular for faithfulness, answer relevancy, context precision/recall
- TruLens
- DeepEval
- LangSmith evaluations
- Custom LLM judge prompts
Rule of thumb
- If you’re tuning retrieval: use Recall@k, MRR, nDCG
- If you’re tuning generation: use faithfulness + correctness
- If you’re measuring product quality: use human/LLM rubric + task success
If you want, I can suggest a specific metric set and evaluation pipeline depending on whether your RAG is for customer support, internal docs, code, or legal/medical.