Prompt
How do I store embeddings and search for nearest neighbors fast?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
To store embeddings and search nearest neighbors fast, use a vector index / vector database rather than plain SQL or brute-force search.
Common options
1) Small scale: in-memory matrix + brute force
If you have up to maybe tens of thousands of vectors, you can:
- store embeddings in an array / table
- compute cosine similarity or dot product against all vectors
- take the top-k
This is simple but gets slow as data grows.
2) Medium to large scale: approximate nearest neighbor (ANN) index
Use a specialized ANN method/library such as:
- FAISS (very common, fast, local library)
- HNSWlib (great recall/speed tradeoff)
- Annoy (simple, read-heavy)
- Vector databases like Pinecone, Weaviate, Milvus, Qdrant, pgvector
These indexes trade a bit of exactness for huge speedups.
Typical storage pattern
Store:
- the vector embedding
- an ID
- metadata you want to filter on
Example schema:
{
"id": "doc_123",
"embedding": [0.12, -0.44, ...],
"text": "optional content",
"metadata": {
"source": "support_docs",
"language": "en"
}
}
How nearest-neighbor search works
- Generate an embedding for your query.
- Search the vector index for the closest stored vectors.
- Return top-k matches.
Similarity metrics:
- Cosine similarity: most common for text embeddings
- Dot product: often used when embeddings are normalized
- Euclidean distance: common for some use cases
If using cosine similarity, normalize vectors first if your library expects it.
Recommended approach by use case
If you want the easiest production setup
Use a vector database:
- Qdrant: easy, open source, strong filtering
- Pinecone: managed service
- Weaviate: feature-rich
- Milvus: scalable
If you want a local library and maximum speed
Use FAISS or HNSWlib.
Example with FAISS:
import faiss
import numpy as np
# embeddings: shape (N, d), dtype float32
embeddings = np.array([...], dtype=np.float32)
index = faiss.IndexFlatIP(embeddings.shape[1]) # inner product
faiss.normalize_L2(embeddings)
index.add(embeddings)
query = np.array([...], dtype=np.float32).reshape(1, -1)
faiss.normalize_L2(query)
scores, ids = index.search(query, k=5)
Performance tips
- Use float32
- Normalize vectors if using cosine similarity
- Batch inserts and queries
- Use ANN indexes for large datasets
- Keep metadata separate from vector storage if needed
- For very large scale, shard or use a managed vector DB
Rule of thumb
- < 50k vectors: brute force may be fine
- 50k–millions: ANN index
- Need filtering + persistence + ops simplicity: vector database
If you want, I can recommend the best tool for your stack (Python/Postgres/JS/etc.) and give you a minimal working example.