Prompt

My ANN recall dropped after adding more vectors, what could be wrong?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

A recall drop after adding more vectors is usually a sign that the ANN index/search parameters or data distribution changed in a way the index can’t keep up with. Common causes:

1) Index too small for the larger dataset

If you added many vectors but kept the same ANN settings, the search may now be too approximate.

  • HNSW: efSearch may be too low
  • IVF / IVFPQ: nprobe may be too low
  • DiskANN / graph-based systems: search beam / hop settings may need to increase
  • FAISS flat vs approximate: if you switched or rebuilt with a more compressed index, recall can drop

Fix: Increase search parameters and/or index size-related parameters:

  • HNSW: raise efSearch (and maybe efConstruction, M)
  • IVF: raise nprobe (and maybe nlist)
  • PQ/compressed indexes: use more bits/subquantizers, or test a non-compressed index

2) Index was not rebuilt / not fully refreshed

If your system supports incremental inserts, the new vectors may be in a “delta” or unoptimized segment, or the main index may not include them well.

Symptoms:

  • New vectors are hard to retrieve
  • Performance changes after inserts
  • Recall drops mainly for queries near newly inserted vectors

Fix: Rebuild/compact/optimize the index, or verify inserts are actually indexed.


3) Dataset grew, but ground truth changed

If you measure recall against an old ground-truth set, adding vectors can make the “top-k true neighbors” change. Your ANN may still be fine, but the benchmark is now stricter.

Check:

  • Recompute exact nearest neighbors on the new full dataset
  • Ensure recall is measured against the updated ground truth

4) Distribution drift from the new vectors

If the added vectors are from a different distribution, cluster structure may shift and old index partitions become less effective.

Examples:

  • New data from a new domain or language
  • Different embedding model/version
  • Different normalization or preprocessing

Fix: Retrain the index quantizer/centroids, rebuild the graph, or separate indices by cohort/domain if appropriate.


5) Vector normalization or metric mismatch

Adding vectors sometimes exposes preprocessing inconsistencies:

  • Some vectors normalized, others not
  • Cosine similarity used but vectors indexed with L2
  • Query and index vectors produced by different model versions

Check:

  • Same embedding model/version?
  • Same normalization applied to all vectors and queries?
  • Same distance metric in training/indexing/search?

6) Capacity/parameter limits exceeded

Some indexes degrade when overfilled:

  • IVF with too few lists for the dataset size
  • HNSW with too-small M
  • PQ with too much compression
  • Shards becoming imbalanced

Fix: Increase index capacity/settings or shard more evenly.


7) Search budget stayed fixed while data increased

Even if the index is unchanged, searching the same number of candidates over a larger corpus usually yields lower recall.

Rule of thumb: More data generally requires more search effort to maintain recall.


Quick debugging checklist

  1. Verify exact recall on the full updated dataset.
  2. Compare before/after with the same queries and same metric.
  3. Increase search parameters:
    • HNSW: efSearch
    • IVF: nprobe
  4. Check vector preprocessing consistency.
  5. Rebuild/retrain the index if data distribution changed.
  6. Inspect whether new vectors are indexed correctly.
  7. Test recall on only the new vectors vs old vectors to isolate the issue.

If you tell me what ANN method/library you’re using (FAISS, HNSWlib, Milvus, Pinecone, Weaviate, etc.) and your index/search parameters, I can suggest the most likely fix.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.