Prompt
I'm building RAG and need to evaluate search vs generation separately
Latest observation
A good way to evaluate a RAG system separately is to split it into two pipelines:
- Retriever / search evaluation — “Did we find the right evidence?”
- Generator evaluation — “Given the evidence, did the model answer well?”
That separation is important because a bad answer may come from:
- bad retrieval,
- bad generation,
- or both.
1) Evaluate search / retrieval
Use a dataset of:
- query
- relevant documents/passages
- optionally graded relevance labels
Common metrics
- Recall@k: Did any relevant chunk appear in the top-k?
- Precision@k: How many of the top-k chunks were relevant?
- MRR: How high was the first relevant result?
- nDCG@k: Better if you have graded relevance
- Hit Rate@k: Whether at least one relevant item was retrieved
What to measure in RAG
Usually the most important retrieval metrics are:
- Recall@k: because if the answer isn’t in retrieved context, generation can’t recover
- MRR / nDCG: because ranking quality matters
Practical note
Make retrieval evaluation at the chunk level, not just document level, because generation usually uses chunks.
2) Evaluate generation separately
To isolate generation, give the model:
- the gold context or
- the retrieved context from an oracle retriever
Then ask:
- does the model answer correctly from that context?
Metrics
Depending on the task:
- Exact Match (EM) for short factual QA
- F1 for span-like answers
- ROUGE / BLEU for summarization-style outputs
- LLM-as-judge for open-ended answers
- Faithfulness / groundedness: is the answer supported by the provided context?
- Answer correctness: is the answer actually right?
Important separate checks
- Answer correctness: right or wrong
- Faithfulness: supported by retrieved context
- Completeness: includes all required points
- Hallucination rate: unsupported claims
3) Best practice: 3-way evaluation
For RAG, it helps to report:
A. Retrieval quality
- Recall@k
- MRR
B. Context quality
- Context precision
- Context recall
- Context relevance
C. Answer quality
- Exact match / F1 / human eval / LLM judge
This gives you a clear diagnosis:
- high retrieval + low generation => generator issue
- low retrieval + high generator on oracle context => retriever issue
- both low => system issue
4) Use oracle vs actual retrieval
A very useful experiment:
Oracle setting
Pass the ground-truth relevant context to the generator.
This measures the generator’s upper bound.
End-to-end setting
Pass the retrieved context.
Compare the two:
- big gap = retrieval bottleneck
- small gap = generator bottleneck or task is hard
5) Recommended evaluation setup
For each query:
- Retrieve top-k chunks
- Score retrieval against labeled relevant chunks
- Generate answer from:
- retrieved chunks
- gold chunks
- Compare generation outputs to reference answers
- Optionally ask an LLM judge for groundedness and correctness
6) Simple metric table
| Component | Metric | Purpose |
|---|---|---|
| Retrieval | Recall@k | Did we retrieve relevant evidence? |
| Retrieval | MRR | How early is the first relevant hit? |
| Retrieval | nDCG@k | Ranking quality |
| Generation | EM/F1 | Answer correctness for QA |
| Generation | ROUGE | Overlap-based quality |
| Generation | Faithfulness | Grounding in context |
| End-to-end | Answer accuracy | Final system quality |
7) If you want a minimal evaluation recipe
If you’re just starting, do this:
- Retriever:
Recall@5,Recall@10,MRR - Generator on gold context:
EM/F1or LLM judge - End-to-end: same answer metric + groundedness
That’s usually enough to pinpoint where the problem is.
If you want, I can also give you:
- a concrete evaluation schema for your dataset,
- a Python evaluation script template, or
- guidance on how to label relevance for chunks.