Prompt
How do I keep vector search latency low at million-vector scale?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
To keep vector search latency low at million-vector scale, focus on four levers:
1) Use an approximate index
Brute-force search over a million vectors is usually too slow. Common ANN (Approximate Nearest Neighbor) options:
- HNSW: very fast low-latency search, good recall, higher memory use
- IVF / IVF-PQ: good for larger datasets, tunable speed/recall tradeoff
- DiskANN / ScaNN / Faiss variants: strong performance depending on workload
If latency is the priority, HNSW is often the first choice for in-memory systems.
2) Reduce the amount of work per query
Tune parameters to inspect fewer candidates:
- HNSW: lower
efSearchfor lower latency, higher for better recall - IVF: lower
nprobe - Use top-k only as large as needed
- Apply metadata filters before or during search if possible
3) Shrink vector and index size
Smaller vectors and indexes mean less memory traffic:
- Use lower-dimensional embeddings if acceptable
- Apply quantization:
- FP16 or INT8 for vectors
- Product Quantization (PQ) for large-scale compression
- Normalize vectors if using cosine similarity so computation is simpler and more predictable
4) Keep the hot path in memory
Latency spikes often come from cache misses or disk access:
- Keep the index RAM-resident
- Avoid page faults and swapping
- Use memory-friendly layouts and vectorized SIMD instructions
- Preload index structures at startup
Practical tuning checklist
If using HNSW
- Increase
Mif recall is too low, but note memory cost grows - Tune
efConstructionfor build quality - Tune
efSearchfor latency/recall balance - Batch queries if your system allows it
If using IVF
- Increase number of centroids for better selectivity
- Tune
nprobecarefully - Consider PQ or OPQ for compressed storage
- Rebalance/retrain the quantizer when data distribution shifts
If using a vector DB
- Ensure ANN index is enabled, not exact search
- Keep frequently queried shards local
- Use proper sharding and replica placement
- Monitor tail latency, not just average latency
System-level tips
- Use batching for throughput, but not if it hurts p95 latency
- Pin CPU cores and avoid noisy neighbors
- Use SIMD and multithreading
- Cache query embeddings if queries repeat
- Co-locate embeddings and metadata where possible
- Measure p95/p99 latency, recall, and memory together
Rule of thumb
For million-scale search:
- Need best latency + enough RAM → HNSW
- Need lower memory + good scale → IVF-PQ or DiskANN-style approaches
- Need exact search → expect higher latency unless heavily optimized and filtered
If you want, I can also give you:
- a Faiss tuning guide,
- a HNSW parameter cheat sheet, or
- a production architecture for low-latency vector search.