Prompt

I'm building RAG and need to evaluate search vs generation separately

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A good way to evaluate a RAG system separately is to split it into two pipelines:

  1. Retriever / search evaluation — “Did we find the right evidence?”
  2. Generator evaluation — “Given the evidence, did the model answer well?”

That separation is important because a bad answer may come from:

  • bad retrieval,
  • bad generation,
  • or both.

1) Evaluate search / retrieval

Use a dataset of:

  • query
  • relevant documents/passages
  • optionally graded relevance labels

Common metrics

  • Recall@k: Did any relevant chunk appear in the top-k?
  • Precision@k: How many of the top-k chunks were relevant?
  • MRR: How high was the first relevant result?
  • nDCG@k: Better if you have graded relevance
  • Hit Rate@k: Whether at least one relevant item was retrieved

What to measure in RAG

Usually the most important retrieval metrics are:

  • Recall@k: because if the answer isn’t in retrieved context, generation can’t recover
  • MRR / nDCG: because ranking quality matters

Practical note

Make retrieval evaluation at the chunk level, not just document level, because generation usually uses chunks.


2) Evaluate generation separately

To isolate generation, give the model:

  • the gold context or
  • the retrieved context from an oracle retriever

Then ask:

  • does the model answer correctly from that context?

Metrics

Depending on the task:

  • Exact Match (EM) for short factual QA
  • F1 for span-like answers
  • ROUGE / BLEU for summarization-style outputs
  • LLM-as-judge for open-ended answers
  • Faithfulness / groundedness: is the answer supported by the provided context?
  • Answer correctness: is the answer actually right?

Important separate checks

  • Answer correctness: right or wrong
  • Faithfulness: supported by retrieved context
  • Completeness: includes all required points
  • Hallucination rate: unsupported claims

3) Best practice: 3-way evaluation

For RAG, it helps to report:

A. Retrieval quality

  • Recall@k
  • MRR

B. Context quality

  • Context precision
  • Context recall
  • Context relevance

C. Answer quality

  • Exact match / F1 / human eval / LLM judge

This gives you a clear diagnosis:

  • high retrieval + low generation => generator issue
  • low retrieval + high generator on oracle context => retriever issue
  • both low => system issue

4) Use oracle vs actual retrieval

A very useful experiment:

Oracle setting

Pass the ground-truth relevant context to the generator.

This measures the generator’s upper bound.

End-to-end setting

Pass the retrieved context.

Compare the two:

  • big gap = retrieval bottleneck
  • small gap = generator bottleneck or task is hard

5) Recommended evaluation setup

For each query:

  1. Retrieve top-k chunks
  2. Score retrieval against labeled relevant chunks
  3. Generate answer from:
    • retrieved chunks
    • gold chunks
  4. Compare generation outputs to reference answers
  5. Optionally ask an LLM judge for groundedness and correctness

6) Simple metric table

ComponentMetricPurpose
RetrievalRecall@kDid we retrieve relevant evidence?
RetrievalMRRHow early is the first relevant hit?
RetrievalnDCG@kRanking quality
GenerationEM/F1Answer correctness for QA
GenerationROUGEOverlap-based quality
GenerationFaithfulnessGrounding in context
End-to-endAnswer accuracyFinal system quality

7) If you want a minimal evaluation recipe

If you’re just starting, do this:

  • Retriever: Recall@5, Recall@10, MRR
  • Generator on gold context: EM/F1 or LLM judge
  • End-to-end: same answer metric + groundedness

That’s usually enough to pinpoint where the problem is.


If you want, I can also give you:

  1. a concrete evaluation schema for your dataset,
  2. a Python evaluation script template, or
  3. guidance on how to label relevance for chunks.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.