Prompt
rag evaluation metrics
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
RAG evaluation metrics usually cover retrieval quality, generation quality, and end-to-end answer quality.
1) Retrieval metrics
These measure whether the retriever finds the right context.
- Recall@k: Did the relevant document appear in the top k results?
- Precision@k: Of the top k retrieved items, how many are relevant?
- MRR (Mean Reciprocal Rank): How early does the first relevant result appear?
- nDCG@k: Rewards relevant results higher in the ranking.
- Hit Rate@k: Whether at least one relevant result is in the top k.
- Context Relevance: How relevant the retrieved chunks are to the query.
2) Generation metrics
These assess how well the LLM answers using the retrieved context.
- Faithfulness / Groundedness: Is the answer supported by the retrieved context?
- Answer Relevance: Does the answer address the user’s question?
- Correctness / Accuracy: Is the answer factually correct?
- Completeness: Does it cover all important parts of the answer?
- Hallucination rate: How often unsupported claims appear.
3) End-to-end RAG metrics
These evaluate the full pipeline from query to final response.
- Exact Match (EM): Does the answer exactly match the reference?
- F1 score: Token overlap with the reference answer.
- BLEU / ROUGE / BERTScore: Similarity to reference responses.
- Task success rate: Whether the system solves the user task.
- Latency: Time taken to retrieve and generate.
- Cost: Token and infrastructure cost per query.
4) Common RAG-specific metrics
These are often used in modern RAG evaluation frameworks:
- Context Precision
- Context Recall
- Context Entity Recall
- Answer Faithfulness
- Answer Relevance
- Context Utilization
5) Practical evaluation approach
A good setup usually combines:
- Retrieval eval on labeled queries
- LLM-based judge for faithfulness/relevance
- Human review for a sample of outputs
- End-to-end task metrics for business success
6) Popular tools
- RAGAS
- TruLens
- LangSmith
- DeepEval
- OpenAI Evals
- Phoenix / Arize
If you want, I can also give you:
- a table of metrics with formulas
- which metrics to use for your use case
- or a Python example using RAGAS.