Prompt
How do I troubleshoot low recall in vector search with embeddings?
Latest observation
Low recall in vector search usually means the right item is in your index somewhere, but your retrieval system isn’t finding it often enough. Troubleshooting works best if you separate the problem into embedding quality, index quality, and query-time retrieval settings.
1) First, verify the embeddings themselves
Low recall can come from the representation, not the ANN index.
Check:
- Are you embedding the right text?
- For documents, are you embedding the full content, title + body, or only a chunk?
- For queries, are you embedding the exact user query or a normalized version?
- Are query and document embeddings from the same model?
- Mixing models often hurts recall badly.
- Are you using the right similarity metric?
- If the model expects cosine similarity but you use raw dot product or L2 without normalization, results may degrade.
- Are vectors normalized when required?
- Many cosine-based setups need normalization before indexing/querying.
- Are you truncating inputs?
- If your text is longer than the model’s token limit, important content may be lost.
- Are chunks too large or too small?
- Too large: embedding becomes “averaged” and less specific.
- Too small: the relevant context may be split across chunks.
Quick test:
Take a handful of known query-document pairs:
- Embed them
- Rank all candidate documents by exact brute-force similarity
- See if the correct document appears near the top
If brute-force recall is also poor, the problem is likely embeddings/chunking/model choice.
2) Compare exact search vs approximate search
If brute-force recall is good but production recall is poor, the ANN index is the likely culprit.
Do this:
- Run the same queries against:
- Exact nearest neighbor search
- Your ANN index
- Compare recall@k
If exact search is good and ANN search is bad, tune the index.
3) Tune ANN index parameters
Different ANN systems fail for different reasons.
HNSW
Common knobs:
- efSearch: increase to improve recall
- efConstruction: increase during indexing for better graph quality
- M / max connections: larger often improves recall, increases memory
Symptoms:
- Low recall at small efSearch
- Large improvement when efSearch is raised
IVF / FAISS
Common knobs:
- nlist: number of centroids
- nprobe: number of clusters searched at query time
Symptoms:
- Low recall if nprobe is too small
- Better recall when nprobe increases
PQ / compression
Product quantization can hurt recall if:
- Compression is too aggressive
- Vector dimensionality is reduced too much
- Training data is not representative
Try:
- More bits per subvector
- More subquantizers
- Less aggressive compression
- A non-compressed baseline for comparison
4) Inspect your distance function and normalization
This is a very common source of silent errors.
Common mistakes:
- Using dot product when you intended cosine
- Forgetting to L2-normalize embeddings
- Mixing normalized and unnormalized vectors
- Using L2 with embeddings trained for angular similarity
Rule of thumb:
- For cosine similarity:
- Normalize vectors
- Then use dot product or cosine
- For L2:
- Use it only if the embedding model/indexing scheme supports it well
5) Check chunking strategy
Low recall often comes from retrieval granularity.
Too few large chunks:
- Each chunk covers too much unrelated text
- Embedding becomes diffuse
Too many tiny chunks:
- Relevant context is split
- Query may match a neighboring chunk instead of the exact one
Good practice:
- Chunk by semantic boundaries if possible
- Add overlap
- Index title/headers with chunks
- Consider hierarchical retrieval:
- retrieve coarse chunks first
- then rerank or expand locally
6) Evaluate with recall metrics, not just qualitative examples
Use a labeled test set if possible.
Measure:
- Recall@k
- MRR
- nDCG
- Hit rate
Create queries with known relevant documents and check:
- Is the relevant doc in top 1, 5, 10, 20?
- Does performance vary by query type?
Also slice by:
- Query length
- Entity-heavy vs semantic queries
- Short vs long documents
- Head vs tail queries
7) Look for data and domain mismatch
If your embeddings were trained on general text but your domain is specialized, recall can be weak.
Examples:
- Legal, biomedical, code, support tickets, product catalogs
Try:
- Domain-specific embedding models
- Fine-tuning
- Query rewriting
- Hybrid search with lexical + vector retrieval
8) Use hybrid retrieval if exact terms matter
Embeddings are great for semantics, but they can miss:
- IDs
- Names
- Rare terms
- Exact phrases
- Numeric constraints
If recall matters, combine:
- BM25 / keyword search
- Vector search
- Then rerank with a cross-encoder or LLM-based reranker
This often improves recall more than just tweaking vector settings.
9) Check filtering and metadata constraints
Sometimes “low recall” is actually over-filtering.
Examples:
- Metadata filters exclude relevant docs
- Language/date/category filters are too strict
- Pre-filtering shrinks the candidate set too much
- Query routing sends requests to the wrong partition/index
Debug by:
- Searching without filters
- Comparing filtered vs unfiltered recall
- Inspecting whether relevant items exist in the filtered subset
10) Validate index freshness and ingestion
If documents are updated or inserted frequently:
- Are all vectors ingested?
- Are there stale embeddings?
- Is the index rebuilt or refreshed correctly?
- Are duplicates or missing records present?
A simple count mismatch between source data and indexed vectors can cause apparent recall loss.
11) Practical debugging workflow
Here’s a good sequence:
- Pick 20–50 known query-relevant pairs
- Test exact brute-force retrieval
- Test ANN retrieval
- If exact is bad → fix embeddings, chunking, model, normalization
- If exact is good but ANN is bad → tune ANN parameters
- If both are okay individually but production is bad → inspect filters, routing, freshness
- If semantic recall is still weak → add hybrid retrieval and reranking
12) Common fixes by symptom
Symptom: recall improves a lot when k increases
- Your relevant item is often present but not ranked high enough
- Improve embeddings, reranking, or query formulation
Symptom: recall improves a lot when efSearch/nprobe increases
- ANN parameters are too aggressive
- Increase search effort
Symptom: exact search is poor too
- Embeddings/chunking/model mismatch
- Try a better model, normalization, or hybrid search
Symptom: only certain query types fail
- Those queries may need lexical matching or special handling
- Add query routing or hybrid retrieval
If you want, I can also give you:
- a step-by-step debugging checklist,
- FAISS/HNSW tuning tips, or
- a Python script to measure recall@k on your dataset.