Prompt

How do I do low-latency similarity search over millions of vectors?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

To do low-latency similarity search over millions of vectors, you usually want an approximate nearest neighbor (ANN) system rather than exact brute-force search.

1) Pick the right index for your latency/recall target

Common choices:

  • HNSW

    • Very popular for low-latency search.
    • Great query speed and strong recall.
    • Higher memory usage.
    • Good default for many applications.
  • IVF / IVF-PQ

    • Good when the dataset is large and memory matters.
    • Faster than brute force, less memory than HNSW.
    • Often used with vector compression.
    • Slightly more tuning.
  • ScaNN / DiskANN / FAISS variants

    • Good at large scale and production use.
    • DiskANN is especially strong when the index doesn’t fit comfortably in RAM.

If you want the simplest high-performance starting point: HNSW.

2) Use embeddings designed for your similarity metric

Make sure your vectors are normalized or represented consistently:

  • Cosine similarity: L2-normalize vectors, then use inner product or cosine distance.
  • Dot product: store raw embeddings and use inner product.
  • Euclidean distance: use L2 directly.

Many systems convert cosine similarity into inner product by normalizing vectors first.

3) Reduce dimensionality or compress if needed

For millions of vectors, memory and cache efficiency matter.

Options:

  • Float32 vectors: highest fidelity, most memory.
  • Float16: half the memory, often acceptable.
  • Quantization: int8 / product quantization (PQ) to reduce memory significantly.
  • PCA / projection: reduce dimension if acceptable for quality.

4) Tune the ANN parameters

For HNSW, key knobs include:

  • M: graph degree; higher = better recall, more memory.
  • efConstruction: higher = better index quality, slower build.
  • efSearch: higher = better recall, higher latency.

For IVF:

  • number of clusters / lists
  • nprobe at query time

Typical approach:

  • start with a target recall,
  • increase search parameter until latency becomes acceptable,
  • then adjust index construction settings.

5) Filter before or during vector search if possible

If you need metadata filters like:

  • tenant ID
  • category
  • timestamp range
  • ACLs

Use:

  • pre-filtering if selective enough,
  • post-filtering if the ANN system supports oversampling,
  • or partitioned indexes by tenant/category.

This can make a big difference in latency.

6) Keep the hot index in RAM

For low latency:

  • keep the ANN index memory-resident,
  • use SSD only for persistence or cold tiers,
  • avoid per-query full scans.

If the dataset grows too large for RAM:

  • compress vectors,
  • shard the index,
  • or use a disk-based ANN like DiskANN.

7) Shard for scale and concurrency

For millions of vectors:

  • split data across shards,
  • query shards in parallel,
  • merge top-k results.

This helps with:

  • throughput,
  • resilience,
  • memory distribution.

8) Measure recall vs latency with a ground-truth set

Don’t optimize blindly. Build a benchmark:

  • exact top-k on a sample dataset,
  • compare ANN results against exact results,
  • track recall@k and p95/p99 latency.

You want to tune for your actual workload, not just raw speed.

9) Practical stack recommendations

Depending on your setup:

  • Fast prototype / production-ready library: FAISS
  • Simple low-latency serving: HNSWlib
  • Managed vector DB: Pinecone, Weaviate, Milvus, Qdrant
  • Very large scale / disk-friendly: DiskANN-based systems

10) A good default recipe

If you’re starting from scratch:

  1. Normalize embeddings if using cosine similarity.
  2. Use HNSW.
  3. Store vectors as float32 initially.
  4. Benchmark recall@10 and p95 latency.
  5. Tune efSearch for latency, M for quality.
  6. Add sharding if the index grows too large.
  7. Add quantization if memory becomes a bottleneck.

If you want, I can also give you:

  • a concrete architecture for 10M / 100M vectors,
  • FAISS/HNSW code examples,
  • or a decision tree for choosing between HNSW, IVF-PQ, and DiskANN.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.