Prompt
Do I need a vector DB if I'm only storing a few thousand embeddings?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
Not necessarily.
If you’re only storing a few thousand embeddings, a full vector database is often overkill. In many cases, a simpler setup is enough:
Good alternatives for small scale
- In-memory search with NumPy / PyTorch if everything fits comfortably in RAM
- FAISS if you want fast similarity search without running a database
- SQLite/Postgres + embeddings in rows if you also need metadata and basic filtering
- Simple brute-force cosine similarity if the dataset is tiny and latency requirements are modest
When a vector DB starts to make sense
Consider a vector DB if you need:
- Millions of vectors
- Frequent inserts/deletes/updates
- Metadata filtering at scale
- Distributed or persistent serving
- Low-latency approximate nearest-neighbor search
- Multi-user / production-grade retrieval infrastructure
Rule of thumb
For a few thousand embeddings, you can usually start with:
- brute-force similarity search, or
- FAISS, or
- Postgres/SQLite if you need persistence and metadata
Then move to a vector DB later if the system grows or operational needs become more complex.
If you want, I can suggest the simplest option based on your stack (Python, JS, Postgres, etc.).