Prompt
I'm building a vector index for millions of records, what architecture should I use
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For millions of records, the best architecture is usually a two-stage retrieval system:
- Candidate retrieval with a vector index
- Optional reranking / filtering with metadata or a smaller model
That gives you good latency, manageable cost, and flexibility.
Recommended architecture
1) Store vectors in a dedicated ANN index
Use an approximate nearest neighbor index such as:
- HNSW — great recall/latency, simple to use, excellent for read-heavy workloads
- IVF-PQ / IVF-Flat — better for very large scale and memory efficiency
- DiskANN — strong for very large datasets, lower memory footprint, good latency
- Managed/vector DB options: Pinecone, Weaviate, Milvus, Qdrant, Vespa, Elasticsearch/OpenSearch vector search
2) Keep metadata in a separate store
Put structured fields in a database or search engine:
- PostgreSQL / MySQL for operational metadata
- Elasticsearch/OpenSearch for keyword + filter-heavy search
- Data warehouse if mostly offline
Then during retrieval:
- vector search returns top K IDs
- metadata store applies filters / joins / business logic
3) Add a reranking layer
For best relevance:
- retrieve top 50–200 candidates from vector search
- rerank with:
- cross-encoder
- LLM reranker
- heuristic scoring + metadata boosts
This is especially useful if queries are ambiguous.
Which index type should you choose?
If you need low latency and fast iteration
HNSW
- Pros: very good recall, simple, great for up to tens of millions depending on RAM
- Cons: memory-heavy
Good choice if:
- data fits mostly in memory
- you want easy operations
- reads dominate writes
If memory is tighter and scale is larger
IVF-PQ
- Pros: compresses vectors, lower memory/storage
- Cons: more tuning, slightly lower recall
Good choice if:
- tens of millions to hundreds of millions
- vectors don’t fit comfortably in RAM
- you can trade some recall for scale
If you expect very large scale and production search
DiskANN
- Pros: designed for large corpora, strong latency/memory tradeoff
- Cons: more operational complexity depending on implementation
Good choice if:
- huge indexes
- you need cost-efficient serving at scale
Practical architecture for millions of records
Option A: Simple and strong default
- Vector DB / ANN service: HNSW-based
- Metadata DB: PostgreSQL
- Object store: raw documents and embeddings history
- Reranker: optional
Best for:
- millions to low tens of millions
- moderate update rates
- easiest production path
Option B: Search-engine hybrid
- OpenSearch/Elasticsearch for keyword + vector + filters
- Postgres for canonical metadata
Best for:
- hybrid search
- faceted filters
- text relevance + vectors together
Option C: Large-scale optimized setup
- DiskANN or IVF-PQ
- Sharded vector service
- Metadata cache + DB
- Reranking service
- Async ingestion pipeline
Best for:
- high QPS
- large indexes
- cost-sensitive infra
Sharding and scaling
For millions of records, plan for:
- sharding by vector ID or tenant
- replication for availability
- separate ingestion and query nodes
- batch index builds / incremental updates
If you have multi-tenant data, decide early whether to:
- isolate by tenant
- shard globally with tenant filters
- use one index per tenant for strong isolation
Key design choices
Vector size
- 384 / 768 dims are common
- higher dims = more memory and slower search
Distance metric
- cosine for embeddings from most modern models
- dot product if model is trained that way
- L2 for some classical use cases
Filtering
If you need strong filters:
- choose an engine with pre-filtering or post-filtering
- beware of filter + ANN interactions causing poor recall
Updates
If records change often:
- prefer systems with good incremental updates
- avoid designs that require frequent full rebuilds unless acceptable
My default recommendation
If you’re starting today and have millions of records:
- Use HNSW if the dataset fits mostly in memory and you want simplicity.
- Use IVF-PQ or DiskANN if memory/cost matters and scale is growing.
- Keep metadata separate.
- Add reranking if search quality matters.
- Build an ingestion pipeline that supports batching, upserts, and periodic reindexing.
Quick decision rule
- < 10M vectors, low-latency, simple ops → HNSW
- 10M–100M+, memory constrained → IVF-PQ or DiskANN
- Need keyword + vector + filters → OpenSearch/Elasticsearch or Vespa
- Need managed, fast deployment → Pinecone/Qdrant/Weaviate/Milvus depending your needs
If you want, I can also suggest a specific architecture based on your constraints:
- vector dimensions
- update rate
- latency target
- QPS
- whether you need filters
- cloud/on-prem preference