Prompt
I'm building a RAG app and need to measure retrieval quality versus answer quality
Latest observation
A good way to evaluate a RAG system is to separate retrieval quality from answer quality. They are related, but not the same:
- Retrieval quality: Did the retriever fetch the right evidence?
- Answer quality: Did the model produce a correct, useful answer, ideally grounded in that evidence?
1) Measure retrieval quality
You want to know whether the relevant chunks/documents are present in the top-k results.
Common retrieval metrics
If you have labeled queries with known relevant docs/chunks:
- Recall@k
Fraction of queries where at least one relevant item appears in the top-k.- Good for RAG because missing evidence is often fatal.
- Precision@k
Fraction of retrieved items in top-k that are relevant.- Useful, but in RAG recall is often more important than precision.
- MRR (Mean Reciprocal Rank)
Rewards placing the first relevant item high in the ranking. - nDCG@k
Better if you have graded relevance (very relevant vs somewhat relevant).
Practical RAG retrieval checks
Also track:
- Context Recall: Did retrieved context contain the answer-supporting facts?
- Context Precision: How much of the retrieved context was actually useful?
- Coverage of gold evidence: Did the retrieval include the specific source snippet(s) needed?
Example retrieval evaluation setup
For each query:
- Annotate one or more “gold” chunks/documents.
- Retrieve top-k.
- Compute:
- Recall@k
- MRR
- nDCG@k
2) Measure answer quality
Answer quality is about whether the final response is correct, complete, and grounded.
Common answer metrics
If you have reference answers:
- Exact Match (EM)
Strict string match; best for short factual answers. - Token-level F1
Useful when wording varies. - Semantic similarity / embedding similarity
Coarser, but can be helpful. - LLM-as-judge scoring
Judge correctness, completeness, clarity, and groundedness on a rubric.
Important RAG-specific answer metrics
These are more meaningful than generic QA metrics:
- Faithfulness / Groundedness
Is the answer supported by the retrieved context? - Answer correctness
Is the answer factually correct? - Completeness
Does it include all required parts? - Citation accuracy
If you cite sources, are citations actually supporting the claims? - Hallucination rate
How often the model invents unsupported facts?
Suggested rubric for answer evaluation
Score each answer 1–5 on:
- Correctness
- Completeness
- Faithfulness to context
- Conciseness
- Helpfulness
3) Best practice: evaluate the pipeline in stages
A RAG app has at least 3 steps:
- Retrieval
- Reranking / context selection
- Generation
Evaluate each separately and together.
Stage-level diagnostics
- Retrieval fails, answer fails → problem is likely retrieval.
- Retrieval succeeds, answer fails → problem is likely generation/prompting.
- Retrieval succeeds, answer is good but unsupported → hallucination risk.
- Retrieval poor but answer still good → maybe the model is using prior knowledge, which may be dangerous in production.
4) Create a test set
You need a labeled eval set of queries.
For each query, store:
- User question
- Gold answer
- Gold supporting passages/chunks
- Optional: difficulty, domain, ambiguity level
Include different query types:
- Fact lookup
- Multi-hop reasoning
- Summarization from documents
- Ambiguous questions
- Out-of-domain / no-answer questions
This helps you see where RAG breaks.
5) Measure no-answer / abstention behavior
If your app should say “I don’t know” when evidence is missing:
- Abstention accuracy
- False answer rate on unanswerable queries
- Calibration of confidence
This matters a lot in RAG because good systems should avoid answering when retrieval is weak.
6) Recommended metric set for production
If you want a practical starter set:
Retrieval
- Recall@5 or Recall@10
- MRR
- nDCG@10
Answer
- Correctness score
- Faithfulness/groundedness score
- Hallucination rate
- Abstention accuracy for unanswerable queries
End-to-end
- Task success rate
- Human satisfaction score
- Citation correctness if using citations
7) Tools you can use
Depending on your stack:
- RAGAS: popular for retrieval/answer faithfulness metrics
- TruLens: evals for groundedness, relevance, feedback
- DeepEval: RAG evaluation workflows
- LangSmith: traces + evals
- LlamaIndex evals: retrieval/response evaluators
8) A simple formula for interpreting results
A useful diagnostic matrix:
| Retrieval | Answer | Likely issue |
|---|---|---|
| Bad | Bad | Retriever/indexing problem |
| Good | Bad | Prompting / generation / context formatting |
| Bad | Good | Risky reliance on parametric knowledge |
| Good | Good | Healthy system |
9) Minimal evaluation workflow
- Build a labeled set of ~50–200 queries.
- For each query, collect gold evidence and gold answer.
- Run retrieval with top-k.
- Score retrieval with Recall@k and MRR.
- Generate answers from retrieved context.
- Score answer correctness and faithfulness.
- Slice results by query type and difficulty.
- Iterate on retrieval, chunking, reranking, prompt, and model.
If you want, I can also give you:
- a concrete evaluation schema for your dataset,
- a Python notebook template to compute retrieval and answer metrics,
- or a metric rubric you can use with LLM-as-judge.