Prompt
vector search latency memory usage
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
Vector search latency and memory usage depend mostly on three things:
- Index type
- Vector dimensionality and count
- Whether you search exact or approximate
1) Latency
Exact search
- Compares the query against every vector
- Latency grows roughly with dataset size
- Good for small datasets, poor for large ones
Approximate nearest neighbor (ANN)
Common indexes:
- HNSW: usually very fast low-latency search, higher memory
- IVF / PQ: lower memory, can be fast, but may reduce recall
- Flat + GPU: fast if hardware is strong, but still scales with data volume
Latency is affected by:
- Number of vectors searched
- Vector dimension
- Recall target
- CPU vs GPU
- Cache locality
- Filtering/metadata constraints
- Concurrency and load
2) Memory usage
Memory is dominated by:
- The raw vectors themselves
- Index overhead
- Metadata
- Replication / sharding overhead
Raw vector size
If you store float32 vectors:
Memory per vector = dimensions × 4 bytes
Examples:
- 768-d vector → 768 × 4 = 3072 bytes ≈ 3 KB
- 1536-d vector → 1536 × 4 = 6144 bytes ≈ 6 KB
So:
- 1 million 768-d float32 vectors ≈ 3 GB just for the vectors
- 1 million 1536-d float32 vectors ≈ 6 GB
Index overhead
- HNSW can add significant overhead because it stores graph links
- IVF adds centroids/lists
- PQ compresses vectors, reducing memory a lot
3) Tradeoff summary
- Lowest latency: HNSW or GPU-based search
- Lowest memory: compressed indexes like PQ/quantization
- Best accuracy: exact search or high-recall ANN
- Best scalability: sharded ANN with compression and filtering
Practical rule of thumb
If you want:
- Fastest queries: use HNSW, tuned for recall/latency
- Lower memory footprint: use IVF-PQ or quantization
- Exact results: use flat search, but expect higher latency
If you want, I can also give:
- a memory sizing formula
- a latency comparison table for HNSW vs IVF vs Flat
- or guidance for a specific setup like Pinecone, FAISS, Milvus, Weaviate, or pgvector