Prompt
How should I compare OpenAI, Cohere, and Voyage AI embeddings for a document search app?
Latest observation
To compare OpenAI, Cohere, and Voyage AI embeddings for a document search app, test them on the thing you actually care about: retrieval quality for your corpus and queries, not just generic benchmarks.
What to compare
1) Retrieval quality
Use a labeled evaluation set if possible:
- Queries: real user questions or search terms
- Relevant docs/passages: the correct matches
- Measure:
- Recall@k: did the right doc appear in the top k?
- MRR: how high was the first relevant result?
- nDCG@k: useful if you have graded relevance
This is the most important comparison.
2) Semantic behavior on your data
Check how each model handles:
- Short queries vs long queries
- FAQ-style queries
- Domain jargon
- Synonyms / paraphrases
- Exact-name lookup vs meaning-based search
3) Cost and latency
For each provider, measure:
- Embedding cost per 1M tokens / documents
- Latency per request
- Throughput / batch performance
- Vector dimensionality and resulting index size
This matters for production search.
4) Chunking sensitivity
Embeddings can perform very differently depending on:
- Chunk size
- Overlap
- Whether you embed titles separately
- Whether queries are short and docs are long
Compare models under the same chunking strategy first.
5) Operational fit
Evaluate:
- API reliability / rate limits
- Batch embedding support
- Version stability
- Data retention / privacy options
- Multilingual support if relevant
A good test plan
Step 1: Build a benchmark set
Create ~100–500 queries from real usage if possible.
For each query, annotate:
- 1–5 relevant documents/chunks
- optionally “highly relevant” vs “somewhat relevant”
Step 2: Embed the same corpus with all three
Keep constant:
- Chunking method
- Vector DB / index type
- Retrieval settings
- Metadata filtering
Step 3: Run identical retrieval tests
For each query:
- Retrieve top 10 or top 20
- Compute Recall@k, MRR, nDCG
Step 4: Inspect failures manually
Look at:
- Missed relevant docs
- Near-duplicate results
- Wrong semantic matches
- Overweighting of boilerplate or headers
Step 5: If you use reranking, test end-to-end
If your app will use a reranker or LLM reranking:
- Compare embeddings alone
- Compare embeddings + reranker Sometimes embedding differences shrink once reranking is added.
Practical comparison criteria
Use a table like this:
| Metric | OpenAI | Cohere | Voyage AI |
|---|---|---|---|
| Recall@10 | |||
| MRR | |||
| nDCG@10 | |||
| Avg latency | |||
| Cost / 1M tokens | |||
| Index size | |||
| Multilingual performance | |||
| Rate limits / stability |
What often matters in practice
For document search, the “best” embedding model is usually the one that:
- Retrieves the right docs consistently on your corpus
- Is affordable at your scale
- Has acceptable latency
- Works well with your chunking and reranking setup
A model that wins on generic benchmarks may lose on your data.
Recommendation
If you want a fast, robust comparison:
- Use real queries
- Measure Recall@k + MRR
- Test with the same chunks and index
- Include cost, latency, and index size
- Validate with a manual review of misses
If you want, I can help you design a small evaluation harness or a benchmark spreadsheet for OpenAI vs Cohere vs Voyage AI.