Prompt

How do I keep vector search latency low at million-vector scale?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

To keep vector search latency low at million-vector scale, focus on four levers:

1) Use an approximate index

Brute-force search over a million vectors is usually too slow. Common ANN (Approximate Nearest Neighbor) options:

  • HNSW: very fast low-latency search, good recall, higher memory use
  • IVF / IVF-PQ: good for larger datasets, tunable speed/recall tradeoff
  • DiskANN / ScaNN / Faiss variants: strong performance depending on workload

If latency is the priority, HNSW is often the first choice for in-memory systems.

2) Reduce the amount of work per query

Tune parameters to inspect fewer candidates:

  • HNSW: lower efSearch for lower latency, higher for better recall
  • IVF: lower nprobe
  • Use top-k only as large as needed
  • Apply metadata filters before or during search if possible

3) Shrink vector and index size

Smaller vectors and indexes mean less memory traffic:

  • Use lower-dimensional embeddings if acceptable
  • Apply quantization:
    • FP16 or INT8 for vectors
    • Product Quantization (PQ) for large-scale compression
  • Normalize vectors if using cosine similarity so computation is simpler and more predictable

4) Keep the hot path in memory

Latency spikes often come from cache misses or disk access:

  • Keep the index RAM-resident
  • Avoid page faults and swapping
  • Use memory-friendly layouts and vectorized SIMD instructions
  • Preload index structures at startup

Practical tuning checklist

If using HNSW

  • Increase M if recall is too low, but note memory cost grows
  • Tune efConstruction for build quality
  • Tune efSearch for latency/recall balance
  • Batch queries if your system allows it

If using IVF

  • Increase number of centroids for better selectivity
  • Tune nprobe carefully
  • Consider PQ or OPQ for compressed storage
  • Rebalance/retrain the quantizer when data distribution shifts

If using a vector DB

  • Ensure ANN index is enabled, not exact search
  • Keep frequently queried shards local
  • Use proper sharding and replica placement
  • Monitor tail latency, not just average latency

System-level tips

  • Use batching for throughput, but not if it hurts p95 latency
  • Pin CPU cores and avoid noisy neighbors
  • Use SIMD and multithreading
  • Cache query embeddings if queries repeat
  • Co-locate embeddings and metadata where possible
  • Measure p95/p99 latency, recall, and memory together

Rule of thumb

For million-scale search:

  • Need best latency + enough RAM → HNSW
  • Need lower memory + good scale → IVF-PQ or DiskANN-style approaches
  • Need exact search → expect higher latency unless heavily optimized and filtered

If you want, I can also give you:

  1. a Faiss tuning guide,
  2. a HNSW parameter cheat sheet, or
  3. a production architecture for low-latency vector search.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.