Prompt

How do I troubleshoot low recall in vector search with embeddings?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Low recall in vector search usually means the right item is in your index somewhere, but your retrieval system isn’t finding it often enough. Troubleshooting works best if you separate the problem into embedding quality, index quality, and query-time retrieval settings.

1) First, verify the embeddings themselves

Low recall can come from the representation, not the ANN index.

Check:

  • Are you embedding the right text?
    • For documents, are you embedding the full content, title + body, or only a chunk?
    • For queries, are you embedding the exact user query or a normalized version?
  • Are query and document embeddings from the same model?
    • Mixing models often hurts recall badly.
  • Are you using the right similarity metric?
    • If the model expects cosine similarity but you use raw dot product or L2 without normalization, results may degrade.
  • Are vectors normalized when required?
    • Many cosine-based setups need normalization before indexing/querying.
  • Are you truncating inputs?
    • If your text is longer than the model’s token limit, important content may be lost.
  • Are chunks too large or too small?
    • Too large: embedding becomes “averaged” and less specific.
    • Too small: the relevant context may be split across chunks.

Quick test:

Take a handful of known query-document pairs:

  • Embed them
  • Rank all candidate documents by exact brute-force similarity
  • See if the correct document appears near the top

If brute-force recall is also poor, the problem is likely embeddings/chunking/model choice.


2) Compare exact search vs approximate search

If brute-force recall is good but production recall is poor, the ANN index is the likely culprit.

Do this:

  • Run the same queries against:
    1. Exact nearest neighbor search
    2. Your ANN index
  • Compare recall@k

If exact search is good and ANN search is bad, tune the index.


3) Tune ANN index parameters

Different ANN systems fail for different reasons.

HNSW

Common knobs:

  • efSearch: increase to improve recall
  • efConstruction: increase during indexing for better graph quality
  • M / max connections: larger often improves recall, increases memory

Symptoms:

  • Low recall at small efSearch
  • Large improvement when efSearch is raised

IVF / FAISS

Common knobs:

  • nlist: number of centroids
  • nprobe: number of clusters searched at query time

Symptoms:

  • Low recall if nprobe is too small
  • Better recall when nprobe increases

PQ / compression

Product quantization can hurt recall if:

  • Compression is too aggressive
  • Vector dimensionality is reduced too much
  • Training data is not representative

Try:

  • More bits per subvector
  • More subquantizers
  • Less aggressive compression
  • A non-compressed baseline for comparison

4) Inspect your distance function and normalization

This is a very common source of silent errors.

Common mistakes:

  • Using dot product when you intended cosine
  • Forgetting to L2-normalize embeddings
  • Mixing normalized and unnormalized vectors
  • Using L2 with embeddings trained for angular similarity

Rule of thumb:

  • For cosine similarity:
    • Normalize vectors
    • Then use dot product or cosine
  • For L2:
    • Use it only if the embedding model/indexing scheme supports it well

5) Check chunking strategy

Low recall often comes from retrieval granularity.

Too few large chunks:

  • Each chunk covers too much unrelated text
  • Embedding becomes diffuse

Too many tiny chunks:

  • Relevant context is split
  • Query may match a neighboring chunk instead of the exact one

Good practice:

  • Chunk by semantic boundaries if possible
  • Add overlap
  • Index title/headers with chunks
  • Consider hierarchical retrieval:
    • retrieve coarse chunks first
    • then rerank or expand locally

6) Evaluate with recall metrics, not just qualitative examples

Use a labeled test set if possible.

Measure:

  • Recall@k
  • MRR
  • nDCG
  • Hit rate

Create queries with known relevant documents and check:

  • Is the relevant doc in top 1, 5, 10, 20?
  • Does performance vary by query type?

Also slice by:

  • Query length
  • Entity-heavy vs semantic queries
  • Short vs long documents
  • Head vs tail queries

7) Look for data and domain mismatch

If your embeddings were trained on general text but your domain is specialized, recall can be weak.

Examples:

  • Legal, biomedical, code, support tickets, product catalogs

Try:

  • Domain-specific embedding models
  • Fine-tuning
  • Query rewriting
  • Hybrid search with lexical + vector retrieval

8) Use hybrid retrieval if exact terms matter

Embeddings are great for semantics, but they can miss:

  • IDs
  • Names
  • Rare terms
  • Exact phrases
  • Numeric constraints

If recall matters, combine:

  • BM25 / keyword search
  • Vector search
  • Then rerank with a cross-encoder or LLM-based reranker

This often improves recall more than just tweaking vector settings.


9) Check filtering and metadata constraints

Sometimes “low recall” is actually over-filtering.

Examples:

  • Metadata filters exclude relevant docs
  • Language/date/category filters are too strict
  • Pre-filtering shrinks the candidate set too much
  • Query routing sends requests to the wrong partition/index

Debug by:

  • Searching without filters
  • Comparing filtered vs unfiltered recall
  • Inspecting whether relevant items exist in the filtered subset

10) Validate index freshness and ingestion

If documents are updated or inserted frequently:

  • Are all vectors ingested?
  • Are there stale embeddings?
  • Is the index rebuilt or refreshed correctly?
  • Are duplicates or missing records present?

A simple count mismatch between source data and indexed vectors can cause apparent recall loss.


11) Practical debugging workflow

Here’s a good sequence:

  1. Pick 20–50 known query-relevant pairs
  2. Test exact brute-force retrieval
  3. Test ANN retrieval
  4. If exact is bad → fix embeddings, chunking, model, normalization
  5. If exact is good but ANN is bad → tune ANN parameters
  6. If both are okay individually but production is bad → inspect filters, routing, freshness
  7. If semantic recall is still weak → add hybrid retrieval and reranking

12) Common fixes by symptom

Symptom: recall improves a lot when k increases

  • Your relevant item is often present but not ranked high enough
  • Improve embeddings, reranking, or query formulation

Symptom: recall improves a lot when efSearch/nprobe increases

  • ANN parameters are too aggressive
  • Increase search effort

Symptom: exact search is poor too

  • Embeddings/chunking/model mismatch
  • Try a better model, normalization, or hybrid search

Symptom: only certain query types fail

  • Those queries may need lexical matching or special handling
  • Add query routing or hybrid retrieval

If you want, I can also give you:

  1. a step-by-step debugging checklist,
  2. FAISS/HNSW tuning tips, or
  3. a Python script to measure recall@k on your dataset.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.