Prompt
I’m unhappy with MongoDB Atlas Vector Search accuracy
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If Atlas Vector Search feels inaccurate, the issue is usually one of these:
-
Embeddings aren’t a good fit
- The model may not capture your domain well.
- Try a stronger embedding model, or one tuned for your content type.
- Make sure query and document embeddings come from the same model.
-
Chunking is off
- Too large: embeddings become vague.
- Too small: you lose context.
- A common fix is smaller, semantically coherent chunks with overlap.
-
You’re relying on vectors alone
- Pure vector search often misses exact terms, IDs, product names, or rare entities.
- Use hybrid retrieval: combine vector search with keyword/metadata filtering.
-
Index configuration is not ideal
- Check similarity metric: cosine vs dot product vs Euclidean.
- Ensure the indexed field is the correct vector field and dimensions match exactly.
-
Top-k / reranking
- Vector search often benefits from retrieving more candidates, then reranking.
- Try larger candidate pools and a reranker or post-filtering step.
-
Metadata filtering
- If you have category, date, language, tenant, or type fields, use them.
- Without filters, the search space can be too broad.
-
Data quality
- Duplicates, noisy text, boilerplate, or OCR errors can hurt relevance.
- Clean the text before embedding.
-
Expectation mismatch
- Vector search is semantic, not exact. It can be “right” semantically but not return the exact item you expected.
Quick debugging checklist
- Verify embedding model consistency
- Inspect a few query/document embedding pairs manually
- Test multiple chunk sizes
- Increase candidate count and rerank
- Add keyword + metadata filtering
- Compare against a baseline BM25/full-text search
If you want, I can help you diagnose this systematically. Send me:
- your embedding model
- chunk size/overlap
- Atlas vector search index config
- a few example queries and the bad results you’re seeing