Prompt
I want a recommendation for evaluating a retrieval-augmented assistant where failures could be in search, reranking, or generation
Latest observation
A good way to evaluate a retrieval-augmented assistant is to separate the pipeline into retrieval, reranking, and generation, then measure each stage with targeted metrics and a small end-to-end set.
Recommended evaluation setup
1) Build a labeled test set
Create a set of user queries with:
- the correct source documents/passages
- the expected answer
- optionally, negative or distractor documents
This lets you tell whether a failure came from:
- search: relevant info never retrieved
- reranking: relevant docs retrieved but not ordered high enough
- generation: relevant context was available but the model answered poorly or hallucinated
2) Measure retrieval quality first
For the search stage, use:
- Recall@k: Did the relevant document appear in the top-k retrieved results?
- MRR: How high was the first relevant result ranked?
- nDCG@k: Useful if you have graded relevance
Interpretation:
- Low Recall@k → search problem
- Good Recall@k but low MRR/nDCG → reranking problem likely
3) Measure reranking quality separately
If you have a candidate set from search and a reranker on top, evaluate:
- Pairwise accuracy: Does the reranker prefer relevant docs over irrelevant ones?
- nDCG@k / MRR on the reranked list
- Top-1 accuracy: Is the best passage actually relevant?
Interpretation:
- Search retrieves relevant docs, but reranker doesn’t surface them near the top → reranker issue
4) Measure generation quality conditioned on gold context
To isolate generation, test the assistant with:
- the gold documents/passages
- or the top relevant retrieved context
Then evaluate:
- Answer correctness vs reference
- Faithfulness / groundedness: Is every claim supported by context?
- Hallucination rate
- Citation accuracy if citations are used
Interpretation:
- If the model fails even with gold context → generation problem
- If it succeeds with gold context but fails in the full pipeline → retrieval/reranking problem
5) Use an error taxonomy
For each failed query, label it as one of:
- Search miss: relevant evidence absent from retrieval
- Ranking miss: evidence retrieved but not prioritized
- Generation error: evidence present but model misused it
- Ambiguous query / bad gold label: evaluation issue
- Knowledge gap in corpus: answer not actually in indexed data
This is very helpful for debugging.
6) Recommended overall methodology
A practical evaluation stack:
- Offline retrieval eval
- Recall@k, MRR, nDCG
- Reranker eval
- nDCG@k, MRR, top-1 accuracy
- Grounded generation eval
- exact match / semantic similarity
- factuality / citation support
- End-to-end task success
- human or rubric-based scoring
- response helpfulness and correctness
7) If you want one simple recommendation
If you only pick a few metrics, I’d recommend:
- Recall@10 for search
- MRR / nDCG@10 for reranking
- Answer correctness + groundedness for generation
- A manual failure audit on a few dozen misses to classify root cause
That combination usually gives the clearest picture of where the system is failing.
If you want, I can also propose a concrete evaluation rubric, a test set template, or a Python notebook-style scoring pipeline.