Prompt
What embedding model should I use for document search, and how do I evaluate retrieval quality?
Latest observation
For document search, the best embedding model depends on your constraints, but a strong default is:
- Best general-purpose choice:
text-embedding-3-large(high quality, widely applicable) - Cheaper/faster choice:
text-embedding-3-small - If you need open-source/on-prem: a strong modern option like bge-large-en-v1.5 or e5-large-v2 (English), or a multilingual variant if needed
How to choose an embedding model
Pick based on these factors:
-
Domain
- General documents: a general-purpose model is fine.
- Legal, medical, finance, code, etc.: test on your own data; domain-specific performance can vary a lot.
-
Language
- English-only vs multilingual matters a lot.
- Use a multilingual embedding model if queries/docs span multiple languages.
-
Latency and cost
- Smaller models are cheaper and faster.
- Larger models often improve recall/semantic matching.
-
Deployment constraints
- SaaS/API models are easiest to start with.
- Open-source models help with privacy, control, and offline deployment.
Good practical recommendation
If you’re starting from scratch:
- Prototype with
text-embedding-3-large - Build an eval set from your real queries
- If cost/latency is too high, compare against
text-embedding-3-smallor an open-source candidate
How to evaluate retrieval quality
You want to measure how well your retrieval system returns the right documents or passages for a query.
1) Create a test set
Build a small labeled dataset:
- Queries: real user questions/searches
- Relevant documents/chunks: which items should be retrieved
- Ideally include:
- Easy queries
- Ambiguous queries
- Rare terminology
- Short and long queries
A good starting point is 50–200 queries with human judgments.
2) Measure standard retrieval metrics
Most common metrics:
- Recall@k
- Of all relevant documents, how many appear in the top k?
- Great for document search where missing relevant docs is costly.
- Precision@k
- Of the top k results, how many are relevant?
- Useful when users inspect only a few results.
- MRR (Mean Reciprocal Rank)
- Rewards placing the first relevant result high.
- nDCG@k
- Good if you have graded relevance levels (highly relevant, somewhat relevant, etc.).
Typical focus for document search:
- Recall@5 / Recall@10
- MRR@10
- nDCG@10
If you use retrieval before an LLM answers questions, Recall@k is often the most important.
3) Evaluate end-to-end if you use RAG
If retrieval feeds a generation model, also measure:
- Answer correctness
- Citation accuracy
- Faithfulness / groundedness
- Whether the retrieved context actually contains the answer
Because a retriever can look good on paper but still fail in the final user experience.
Practical evaluation workflow
Offline evaluation
- Collect real queries from logs or create representative ones
- Label relevant docs/chunks
- Run retrieval with candidate models
- Compare metrics:
- Recall@k
- MRR@k
- nDCG@k
- Inspect failures manually
Human review
For a sample of queries, check:
- Are retrieved passages actually useful?
- Are there near-duplicate results?
- Does chunking hurt or help?
- Are false positives semantically related but not answer-bearing?
Online evaluation
Once deployed:
- Track click-through rate
- Measure query reformulation rate
- Measure time-to-success
- A/B test embedding models or retrieval strategies
Tips to improve retrieval beyond the embedding model
Even a great embedding model can underperform if the retrieval pipeline is weak.
Important levers:
- Chunking strategy
- Too large: noisy embeddings
- Too small: missing context
- Metadata filtering
- Restrict by source, date, product, language, permissions
- Hybrid search
- Combine vector search + keyword/BM25
- Often improves exact-match queries and rare terms
- Reranking
- Use a cross-encoder or reranker on top results
- Often one of the biggest quality boosts
- Query rewriting
- Expand abbreviations, resolve pronouns, normalize terminology
A simple decision rule
- If you want best quality quickly: use
text-embedding-3-large - If you need lower cost: try
text-embedding-3-small - If you need open-source: benchmark bge or e5 models
- Always choose the model by measuring Recall@k / MRR / nDCG on your own queries
If you want, I can also give you:
- a step-by-step evaluation template, or
- a Python script to benchmark two embedding models on your dataset.