Prompt

How do I debug poor recall with approximate nearest neighbor search?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

To debug poor recall in approximate nearest neighbor (ANN) search, work from the query path outward and compare against an exact baseline. Here’s a practical checklist.

1) First verify the definition of recall

Make sure you’re measuring the right thing:

  • Recall@k = fraction of the exact top-k neighbors that appear in the ANN top-k results.
  • For a query, if exact top-10 is {a,b,c,...} and ANN returns 7 of them, recall@10 = 0.7.
  • Confirm whether you’re evaluating:
    • per-query recall,
    • average recall over queries,
    • macro vs micro averaging,
    • filtered vs unfiltered search.

A common mistake is comparing different metrics or using inconsistent k.


2) Build an exact baseline

Before tuning ANN, get a trusted ground truth:

  • Run brute-force search on a small sample of queries and/or a smaller subset of the corpus.
  • Use the same embedding vectors, same normalization, and same distance metric.
  • Check that exact search itself is correct.

If exact and ANN disagree unexpectedly, the issue is often upstream data or metric handling, not the ANN index.


3) Check preprocessing and distance consistency

Poor recall often comes from a mismatch between how vectors were indexed and how queries are formed.

Common pitfalls

  • Normalization mismatch
    • Cosine similarity usually requires L2-normalized vectors.
    • If the index uses dot product but data/queries are normalized inconsistently, recall can collapse.
  • Wrong metric
    • Euclidean vs cosine vs inner product mismatch.
  • Data transformations
    • PCA, whitening, quantization, truncation, or mean-centering applied to one side but not the other.
  • NaNs / infs / zero vectors
    • These can silently degrade results.

Sanity checks

  • Confirm vector norms.
  • Compare a few manual distances for known pairs.
  • Verify query preprocessing matches ingestion preprocessing exactly.

4) Inspect index build quality

ANN recall depends heavily on how the index was constructed.

For graph-based indexes (e.g., HNSW, NSG)

  • Increase graph connectivity:
    • M / out-degree
  • Increase construction effort:
    • efConstruction
  • Make sure the graph is not too sparse or poorly connected.
  • Check whether insertion order or sharding is hurting graph quality.

For IVF / clustering-based indexes

  • Increase number of clusters / lists:
    • too few clusters => coarse partitions, low recall
  • Increase probes at query time:
    • nprobe / number of clusters searched
  • Inspect cluster balance:
    • highly skewed clusters can hurt recall.

For PQ / compressed indexes

  • Compression can reduce recall.
  • Check:
    • codebook size,
    • number of subquantizers,
    • residual quantization settings,
    • whether re-ranking with original vectors is enabled.

A good debugging tactic: start with a less-compressed or less-approximate configuration and gradually tighten approximation until recall drops.


5) Tune query-time search parameters

Low recall is often simply under-searching.

Graph indexes

  • Increase:
    • efSearch
    • candidate queue size
  • If recall improves dramatically with higher efSearch, the issue is query budget, not index construction.

IVF / inverted file indexes

  • Increase:
    • nprobe
    • number of explored lists

Tree / beam methods

  • Increase:
    • beam width,
    • backtracking/search budget.

Early termination

  • Check whether a hard time or distance threshold is stopping search too early.

Plot recall vs latency as you increase the search budget; this often reveals whether the system is just under-provisioned.


6) Compare ANN results to exact results for failing queries

Focus on the worst queries:

  • Identify queries with very low recall.
  • For each, compare:
    • exact neighbors,
    • ANN neighbors,
    • candidate set returned internally if possible.

Look for patterns:

  • all missed neighbors are from a specific region,
  • queries with large norms,
  • queries near cluster boundaries,
  • queries with many near-ties,
  • queries involving rare vectors.

This can point to cluster imbalance, quantization error, or graph connectivity problems.


7) Check dataset characteristics

ANN quality depends on data distribution.

  • Dimensionality
    • High-dimensional spaces are harder to index well.
  • Intrinsic dimensionality
    • If data lies on complex manifolds, graph or PQ behavior may vary.
  • Skew / hubs
    • Some points may appear in many neighborhoods.
  • Duplicates / near-duplicates
    • Can make exact ranking unstable and hurt measured recall if ties are handled differently.
  • Class imbalance / sharding imbalance
    • If data is partitioned unevenly, some shards may be undersearched.

8) Verify the candidate reranking step

If your system uses a two-stage pipeline:

  1. ANN retrieves candidates
  2. Exact distances rerank them

Then poor recall can come from either stage.

  • If candidate generation misses the true neighbors, increase ANN budget.
  • If reranking is wrong:
    • you may be using the wrong vectors,
    • not enough candidates,
    • or the reranker metric differs from the index metric.

Check whether the ANN stage returns enough candidates for reranking, e.g. top_k * oversample_factor.


9) Instrument the search

Add diagnostics:

  • number of nodes visited,
  • candidate queue sizes,
  • clusters probed,
  • distance computations,
  • pruning reasons,
  • fraction of searches hitting max budget,
  • query latency distribution.

If many searches hit the budget ceiling, recall is likely limited by search effort.


10) Use controlled experiments

Change one variable at a time.

Examples:

  • exact search vs ANN on 1k vectors,
  • same index, higher efSearch,
  • same search budget, higher M,
  • same budget, different normalization.

Track:

  • recall@k,
  • latency,
  • memory,
  • build time.

This helps separate:

  • data problems,
  • index construction problems,
  • query-time parameter problems.

11) Common failure modes and likely fixes

  • Recall is uniformly low
    • metric mismatch, normalization issue, or too aggressive compression.
  • Recall improves a lot when search budget increases
    • query-time parameters too low.
  • Recall is poor only for some queries
    • cluster imbalance, graph connectivity, skewed data, boundary queries.
  • Recall is good on small samples but bad in production
    • sharding, filtering, dynamic insertions, or preprocessing drift.
  • Recall dropped after adding quantization
    • need more codebooks, larger codes, or reranking.

12) A simple debugging sequence

If you want a quick path:

  1. Validate metric + normalization.
  2. Compare ANN against exact on a small sample.
  3. Increase search budget dramatically.
  4. If recall improves, tune query-time parameters.
  5. If not, rebuild index with stronger construction settings.
  6. If compression is used, test without compression or with reranking.
  7. Inspect worst-case queries and data skew.

If you tell me which ANN method you’re using—HNSW, IVF, FAISS PQ, ScaNN, Annoy, DiskANN, etc.—I can give a more specific debugging playbook and the key parameters to inspect.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.