Prompt
What should I use to compare embedding models on downstream tasks?
Latest observation
To compare embedding models on downstream tasks, the best choice is usually:
1) Use a task-specific benchmark
Pick evaluation sets that match your real use case:
- Retrieval / semantic search: recall@k, MRR, nDCG
- Clustering: silhouette score, adjusted mutual information, purity
- Classification: accuracy, F1, ROC-AUC
- Semantic similarity / STS: Pearson/Spearman correlation
- Reranking / matching: MAP, MRR, NDCG
If you care about practical performance, evaluate embeddings by training or plugging them into the downstream task rather than relying only on intrinsic similarity tests.
2) For a broad, standardized comparison, use MTEB
The Massive Text Embedding Benchmark (MTEB) is the most common general benchmark for text embedding models. It covers many tasks:
- classification
- clustering
- reranking
- retrieval
- STS
- summarization-related tasks
It’s a good default if you want an overall comparison across multiple task types.
3) Also check BEIR for retrieval-focused evaluation
If your main use case is search or retrieval, BEIR is a strong benchmark suite. It’s widely used for measuring generalization across diverse retrieval datasets.
4) Use a small internal eval set
Benchmarks are useful, but the most reliable comparison is often on:
- your own documents
- your own queries
- your own labels / judgments
That tells you which embedding model actually works best for your data distribution.
Practical recommendation
- If you want a single broad benchmark: use MTEB
- If you want retrieval/search specifically: use BEIR
- If you want the best real-world decision: create an internal eval set and test on the downstream task directly
If you want, I can also suggest a concrete evaluation setup based on your task (search, clustering, classification, RAG, etc.).