Prompt
I'm unhappy with OpenSearch vector search recall
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
Sorry to hear that — low recall with vector search in OpenSearch is usually fixable, but the cause can be a few different things.
The main levers are:
-
Index-time ANN settings
m,ef_construction(for HNSW)- vector quantization / compression settings, if enabled
-
Query-time settings
knum_candidates/ef_searchequivalent- whether you’re filtering too aggressively before ANN has enough candidates
-
Data quality
- vector normalization
- embedding model mismatch
- duplicate or near-duplicate items
- dimension or distance-metric mismatch
-
Shard and routing effects
- too many shards
- routing causing candidate starvation
- uneven shard distribution
-
Ground truth expectations
- ANN recall vs exact nearest neighbors
- metric choice: cosine vs dot product vs L2
If you want, I can help you troubleshoot systematically. The fastest path is for you to share:
- OpenSearch version
- index mapping for the vector field
- how you indexed the vectors
- the exact search request
- approximate corpus size
- vector dimension and distance metric
- whether you use filters
- what “bad recall” means in your case
If you paste those, I can tell you likely bottlenecks and the specific settings to change.
Brands