Prompt

What embedding model should I use for document search, and how do I evaluate retrieval quality?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

For document search, the best embedding model depends on your constraints, but a strong default is:

  • Best general-purpose choice: text-embedding-3-large (high quality, widely applicable)
  • Cheaper/faster choice: text-embedding-3-small
  • If you need open-source/on-prem: a strong modern option like bge-large-en-v1.5 or e5-large-v2 (English), or a multilingual variant if needed

How to choose an embedding model

Pick based on these factors:

  1. Domain

    • General documents: a general-purpose model is fine.
    • Legal, medical, finance, code, etc.: test on your own data; domain-specific performance can vary a lot.
  2. Language

    • English-only vs multilingual matters a lot.
    • Use a multilingual embedding model if queries/docs span multiple languages.
  3. Latency and cost

    • Smaller models are cheaper and faster.
    • Larger models often improve recall/semantic matching.
  4. Deployment constraints

    • SaaS/API models are easiest to start with.
    • Open-source models help with privacy, control, and offline deployment.

Good practical recommendation

If you’re starting from scratch:

  • Prototype with text-embedding-3-large
  • Build an eval set from your real queries
  • If cost/latency is too high, compare against text-embedding-3-small or an open-source candidate

How to evaluate retrieval quality

You want to measure how well your retrieval system returns the right documents or passages for a query.

1) Create a test set

Build a small labeled dataset:

  • Queries: real user questions/searches
  • Relevant documents/chunks: which items should be retrieved
  • Ideally include:
    • Easy queries
    • Ambiguous queries
    • Rare terminology
    • Short and long queries

A good starting point is 50–200 queries with human judgments.

2) Measure standard retrieval metrics

Most common metrics:

  • Recall@k
    • Of all relevant documents, how many appear in the top k?
    • Great for document search where missing relevant docs is costly.
  • Precision@k
    • Of the top k results, how many are relevant?
    • Useful when users inspect only a few results.
  • MRR (Mean Reciprocal Rank)
    • Rewards placing the first relevant result high.
  • nDCG@k
    • Good if you have graded relevance levels (highly relevant, somewhat relevant, etc.).

Typical focus for document search:

  • Recall@5 / Recall@10
  • MRR@10
  • nDCG@10

If you use retrieval before an LLM answers questions, Recall@k is often the most important.

3) Evaluate end-to-end if you use RAG

If retrieval feeds a generation model, also measure:

  • Answer correctness
  • Citation accuracy
  • Faithfulness / groundedness
  • Whether the retrieved context actually contains the answer

Because a retriever can look good on paper but still fail in the final user experience.


Practical evaluation workflow

Offline evaluation

  1. Collect real queries from logs or create representative ones
  2. Label relevant docs/chunks
  3. Run retrieval with candidate models
  4. Compare metrics:
    • Recall@k
    • MRR@k
    • nDCG@k
  5. Inspect failures manually

Human review

For a sample of queries, check:

  • Are retrieved passages actually useful?
  • Are there near-duplicate results?
  • Does chunking hurt or help?
  • Are false positives semantically related but not answer-bearing?

Online evaluation

Once deployed:

  • Track click-through rate
  • Measure query reformulation rate
  • Measure time-to-success
  • A/B test embedding models or retrieval strategies

Tips to improve retrieval beyond the embedding model

Even a great embedding model can underperform if the retrieval pipeline is weak.

Important levers:

  • Chunking strategy
    • Too large: noisy embeddings
    • Too small: missing context
  • Metadata filtering
    • Restrict by source, date, product, language, permissions
  • Hybrid search
    • Combine vector search + keyword/BM25
    • Often improves exact-match queries and rare terms
  • Reranking
    • Use a cross-encoder or reranker on top results
    • Often one of the biggest quality boosts
  • Query rewriting
    • Expand abbreviations, resolve pronouns, normalize terminology

A simple decision rule

  • If you want best quality quickly: use text-embedding-3-large
  • If you need lower cost: try text-embedding-3-small
  • If you need open-source: benchmark bge or e5 models
  • Always choose the model by measuring Recall@k / MRR / nDCG on your own queries

If you want, I can also give you:

  1. a step-by-step evaluation template, or
  2. a Python script to benchmark two embedding models on your dataset.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.