Prompt
what vector database should I use for embeddings and metadata filtering?
Latest observation
It depends on your scale, latency needs, and how much filtering you need, but here’s the practical shortlist:
Best general-purpose picks
1. Postgres + pgvector
Best if: you want simplicity, strong metadata filtering, and your scale is moderate.
- Great for:
- embeddings
- structured metadata filters
- transactions
- easy operations if you already use Postgres
- Pros:
- one database for vectors + metadata
- SQL filtering is very flexible
- easy to develop and deploy
- Cons:
- not ideal for very large-scale ANN workloads
- performance may lag specialized vector DBs at high volume
Use this if you’re building an MVP, internal tool, or app with up to millions of vectors and need robust filtering.
2. Pinecone
Best if: you want managed infrastructure and production-grade vector search with filtering.
- Great for:
- low operational overhead
- fast vector search
- metadata filtering
- Pros:
- fully managed
- scales well
- easy API
- Cons:
- can be expensive
- less flexible than SQL-backed approaches
- vendor lock-in
Use this if you want to move fast and don’t want to manage infrastructure.
3. Weaviate
Best if: you want a feature-rich vector database with good metadata support.
- Great for:
- hybrid search
- schema-based data
- filtering
- Pros:
- strong feature set
- supports hybrid keyword + vector search
- open source + managed options
- Cons:
- more moving parts than Postgres
- operational complexity if self-hosted
Use this if you need more advanced search features and want an open-source option.
4. Qdrant
Best if: metadata filtering is important and you want a strong vector-native system.
- Great for:
- filtering
- payload-based metadata
- fast ANN search
- Pros:
- excellent filter performance
- simple API
- open source and managed offerings
- Cons:
- not as natural as SQL for complex relational queries
- smaller ecosystem than Postgres
Use this if you want a dedicated vector DB with good filtering and solid performance.
Quick recommendation by use case
- Small/medium app, strong metadata filtering, easiest stack:
Postgres + pgvector - Managed production system, minimal ops:
Pinecone - Open-source vector DB with strong filtering:
Qdrant - Hybrid search + richer vector features:
Weaviate
My default recommendation
If you’re unsure: start with Postgres + pgvector unless you already know you’ll need very high-scale vector search or low-latency managed indexing. It’s usually the best combination of:
- simplicity
- filtering power
- cost
- portability
If you want, I can also recommend one based on your:
- expected number of vectors
- filter complexity
- latency target
- budget
- cloud/on-prem preference