Prompt

need nearest-neighbor search for millions of vectors

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

For millions of vectors, you usually want an approximate nearest-neighbor (ANN) index, not a brute-force scan.

Best practical options

1) FAISS

  • Best choice if you care about speed + scale
  • Supports CPU and GPU
  • Good index types:
    • HNSW: great recall/speed, simple to use
    • IVF / IVFPQ: better for very large datasets, memory-efficient
    • Flat: exact search, only feasible for smaller sets or GPU

2) hnswlib

  • Very easy to use
  • Excellent for high recall and fast queries
  • Good for millions of vectors if memory is acceptable
  • No GPU support

3) ScaNN (Google)

  • Strong performance on embeddings
  • Good for large-scale retrieval
  • More opinionated / less general than FAISS

4) Vector databases

If you also need persistence, filtering, metadata, and production ops:

  • Pinecone
  • Milvus
  • Weaviate
  • Qdrant
  • pgvector for smaller scale or simpler setups

What to choose

  • Fastest development: hnswlib
  • Most flexible / scalable: FAISS
  • Production DB + metadata: Qdrant, Milvus, Pinecone
  • Exact search: only if your dataset is small enough or you can use GPU

Rule of thumb

  • 1M–50M vectors: HNSW or FAISS IVF/HNSW
  • 50M+ vectors: FAISS IVF-PQ / disk-backed approaches / vector DB
  • If vectors are high-dimensional and memory-heavy, use compression like PQ

Example: FAISS HNSW

import faiss
import numpy as np

d = 768
xb = np.random.random((1_000_000, d)).astype('float32')
xq = np.random.random((10, d)).astype('float32')

index = faiss.IndexHNSWFlat(d, 32)  # 32 neighbors in graph
index.hnsw.efConstruction = 200
index.add(xb)

index.hnsw.efSearch = 64
D, I = index.search(xq, 10)  # top-10 nearest neighbors

Example: hnswlib

import hnswlib
import numpy as np

dim = 768
num_elements = 1_000_000

data = np.random.random((num_elements, dim)).astype(np.float32)
queries = np.random.random((10, dim)).astype(np.float32)

index = hnswlib.Index(space='l2', dim=dim)
index.init_index(max_elements=num_elements, ef_construction=200, M=16)
index.add_items(data)
index.set_ef(64)

labels, distances = index.knn_query(queries, k=10)

Important considerations

  • Distance metric: cosine, L2, or inner product?
  • Recall vs latency: higher recall usually means slower search
  • Memory budget: HNSW can use a lot of RAM
  • Update frequency: some indexes handle inserts better than others
  • Batching: query in batches for much better throughput

If you tell me:

  1. vector dimension
  2. number of vectors
  3. metric (cosine/L2/dot)
  4. latency target
  5. memory budget
    I can recommend the best index and settings.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.