Prompt
I'm building a recommendation engine with embeddings, any database suggestions?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
Yes — for an embedding-based recommendation engine, the best database depends on your scale, latency needs, and whether you want a single system or a specialized vector database.
Good options
1. PostgreSQL + pgvector
Best if you want simplicity and already use Postgres.
Pros
- Easy to adopt if you already have relational data
- Supports metadata filtering alongside vector search
- Good for prototypes, MVPs, and moderate scale
- Operationally simpler than adding a new DB
Cons
- Not ideal for very large-scale, high-QPS vector search
- Fewer vector-specific optimizations than dedicated vector DBs
Use if
- You want one database for users/items/embeddings
- Your dataset is up to a few million vectors, depending on latency requirements
2. Pinecone
Best managed vector DB for production-grade similarity search.
Pros
- Fully managed
- Low-latency ANN search
- Easy scaling and indexing
- Good metadata filtering
Cons
- Can be expensive at scale
- Vendor lock-in
- Not a general-purpose relational DB
Use if
- You want minimal ops and strong vector-search performance
3. Weaviate
Good open-source / managed hybrid option.
Pros
- Vector search + metadata filtering
- Hybrid search support
- Flexible schema
- Can self-host or use managed
Cons
- More moving parts than Postgres
- Requires tuning/ops knowledge for self-hosting
Use if
- You want a vector-native system but also flexibility
4. Milvus
Strong choice for large-scale vector search.
Pros
- Designed for large vector workloads
- Scales well
- Open source
- Good ANN performance
Cons
- More operational complexity
- Usually paired with other systems for metadata/transactions
Use if
- You expect large catalogs and high throughput
5. OpenSearch / Elasticsearch
Useful if you want search + vectors together.
Pros
- Great for combining keyword and vector search
- Good filtering and ranking pipelines
- Familiar for search teams
Cons
- Vector search is not as specialized as dedicated vector DBs
- Can be heavy/costly
Use if
- Recommendations are part of a broader search/discovery platform
Quick recommendation by stage
- Prototype / MVP: PostgreSQL + pgvector
- Managed production with low ops: Pinecone
- Open-source vector-native: Weaviate or Milvus
- Search + recommendations together: OpenSearch/Elasticsearch
What matters most for your choice
- Dataset size — number of items/vectors
- Latency target — sub-50ms? sub-200ms?
- Update frequency — static catalog vs frequent writes
- Filtering needs — category, locale, price, availability
- Budget / ops tolerance — managed vs self-hosted
- Hybrid search — keyword + vector ranking
Practical architecture
For many recommendation systems:
- Store core app data in PostgreSQL
- Store embeddings in pgvector or a dedicated vector DB
- Add Redis for caching hot recommendations
- Use a feature store / event pipeline if you need real-time personalization
If you want, I can also suggest the best database for your specific case if you tell me:
- number of items
- number of users
- expected QPS
- whether you need filters
- cloud/on-prem preference