Prompt
How do I compare embedding models for retrieval quality?
Latest observation
To compare embedding models for retrieval quality, you want to evaluate them on the task they’ll actually do: given a query, rank relevant documents highly.
Here’s a practical framework.
1) Define the retrieval task
Be explicit about:
- Query type: short search queries, questions, semantic lookup, code search, etc.
- Document type: passages, chunks, product descriptions, support articles, code snippets.
- Language/domain: general, biomedical, legal, multilingual, etc.
- Retrieval setup: top-k retrieval over a corpus, hybrid search, reranking afterward, etc.
A model that’s great for semantic similarity may not be best for your specific retrieval domain.
2) Build or choose a labeled evaluation set
You need query–document relevance labels.
Best option: human-labeled data
For each query, label candidate docs as:
- relevant
- partially relevant
- not relevant
Even a few hundred queries can be enough for comparison if the labels are decent.
If you don’t have labels
Use:
- existing benchmark datasets
- implicit feedback: clicks, dwell time, conversions
- synthetic labels from strong LLM-based judgment, then spot-check
Common retrieval benchmarks:
- BEIR for diverse zero-shot IR
- MTEB retrieval tasks
- MS MARCO for web-style passage retrieval
- domain-specific datasets if you have them
3) Use ranking metrics, not just similarity scores
For retrieval, compare how well each embedding model ranks relevant items.
Core metrics
- Recall@k: did any relevant doc appear in the top-k?
- Precision@k: among top-k, how many are relevant?
- MRR@k: rewards ranking the first relevant item highly
- nDCG@k: best when relevance has graded labels
- MAP: useful when there are multiple relevant docs per query
Common choices
- If you only care whether at least one relevant result is found: Recall@k
- If top result quality matters: MRR@k
- If relevance is graded: nDCG@k
For embedding retrieval, a standard report is:
- Recall@1, Recall@5, Recall@10
- MRR@10
- nDCG@10
4) Keep the retrieval pipeline identical
When comparing models, hold everything else constant:
- same corpus/chunking
- same preprocessing
- same vector database/index settings
- same similarity metric
- same query text
- same top-k
Otherwise you won’t know whether improvements come from the embeddings or the pipeline.
Important:
- If one model uses cosine similarity and another works better with dot product, be consistent with the model’s recommended normalization/score setup, but compare fairly.
- Use the same chunking strategy across models.
5) Evaluate on your real corpus
Generic benchmarks are useful, but the best test is your own data.
A model that scores well on public benchmarks may underperform on:
- your domain vocabulary
- your document style
- your query intent
- your chunk sizes
If possible, create a small gold set from your own content:
- 50–200 queries
- 5–20 judged documents per query
That often reveals the winner quickly.
6) Compare against a lexical baseline
Always include something like:
- BM25
- or another keyword baseline
Why:
- It tells you whether embeddings add value at all.
- Sometimes BM25 beats embeddings on exact-match or entity-heavy queries.
A strong retrieval system often combines:
- BM25 + embeddings + reranker
7) Inspect failure cases
Metrics alone don’t tell the whole story. Look at examples where:
- relevant docs are missed
- semantically similar but incorrect docs are retrieved
- top results are too broad or too narrow
This helps you understand whether the model is bad at:
- lexical matching
- domain terminology
- entity resolution
- query intent
- long-document chunking
8) Check latency, memory, and cost
Retrieval quality is only part of the tradeoff. Also compare:
- embedding dimensionality
- indexing memory
- encoding speed
- query latency
- cost per 1M tokens/documents
A slightly better model may not be worth 3× the cost.
9) Use statistical significance if possible
If two models are close, measure whether the difference is meaningful:
- bootstrap confidence intervals
- paired significance tests over queries
This matters when improvements are small.
10) A simple evaluation recipe
If you want a quick, solid comparison:
- Build a query set with relevance labels.
- Embed the same corpus with each model.
- Retrieve top-10 for each query.
- Compute:
- Recall@10
- MRR@10
- nDCG@10
- Compare to BM25.
- Review 20–30 error cases.
- Choose the best balance of quality, speed, and cost.
Practical advice
If you’re early-stage
Use:
- your real corpus
- 50–100 representative queries
- Recall@10 and MRR@10
- BM25 baseline
If you’re more mature
Add:
- graded judgments
- reranker evaluation
- latency/cost benchmarks
- confidence intervals
Common pitfalls
- Comparing models on different chunking strategies
- Using only cosine similarity scores without retrieval metrics
- Evaluating on queries too similar to the training data
- Overfitting to a small test set
- Ignoring lexical baselines
- Measuring only embedding similarity instead of end-to-end retrieval
If you want, I can also give you:
- a Python evaluation script for comparing embedding models, or
- a checklist for building a retrieval benchmark from your own data.