Prompt
What database should I use for millions of vectors with low latency?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
For millions of vectors with low latency, the best choice is usually a dedicated vector database rather than a general-purpose database with a vector extension.
Good options
- Pinecone – very popular for low-latency vector search at scale, managed, easy to operate.
- Weaviate – solid open-source/managed option, good filtering + hybrid search.
- Milvus – strong for large-scale vector workloads, open-source, flexible, more ops-heavy.
- Qdrant – lightweight, fast, great filtering, simpler to run than Milvus.
- FAISS – extremely fast library, but not a full database; best if you’re building your own service.
Quick recommendation
- If you want easiest production setup: Pinecone
- If you want open source + good performance: Qdrant or Weaviate
- If you need very large scale and control: Milvus
- If you’re embedding inside an app and can manage infra yourself: FAISS
What to look for
For millions of vectors and low latency, make sure the DB supports:
- Approximate nearest neighbor (ANN) search
- Filtering by metadata
- Index types like HNSW, IVF, or PQ
- Fast upserts
- Horizontal scaling / sharding
- Persistence and backups
Important note
If “millions” means maybe 1–50 million vectors, Qdrant, Weaviate, Pinecone, or Milvus are all reasonable.
If you need sub-10 ms latency at high QPS, the implementation and index tuning matter as much as the database choice.
If you want, I can recommend one based on your:
- vector dimension,
- QPS,
- update frequency,
- cloud/on-prem preference,
- and whether you need metadata filtering.