Prompt
What database should I use for embeddings with metadata filters and tenant access?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
If you need embeddings + metadata filters + tenant access control, the best choice depends on your scale and how much infra you want to operate. Here’s the practical shortlist:
Best default: PostgreSQL + pgvector
Use this if:
- You already use Postgres
- You want metadata filtering and tenant isolation in one system
- You have moderate-to-large scale, not billion-scale vectors
- You want the simplest operational model
Why it works well:
pgvectorstores embeddings directly in Postgres- Metadata filtering is easy with SQL (
WHERE tenant_id = ... AND status = ...) - Access control can be enforced with:
- app-level tenant checks
- Row Level Security (RLS)
- separate schemas or databases per tenant if needed
- Great for hybrid retrieval: vector + relational filters
Typical schema:
tenant_idembedding vector(...)- metadata columns like
doc_type,created_at,source,tags - optional JSONB for flexible metadata
Best managed vector DB: Pinecone
Use this if:
- You want a fully managed service
- You need high scale and low ops
- You want vector search + metadata filters out of the box
Pros:
- Strong filtering support
- Easy scaling
- Good production reliability
Cons:
- Less natural for strict relational tenant logic than Postgres
- More vendor lock-in
- Typically more expensive than Postgres-based setups
Best if you already use cloud analytics infra: MongoDB Atlas Vector Search
Use this if:
- Your app already lives in MongoDB
- Metadata is flexible/document-like
- You want combined document + vector queries
Good for:
- JSON-heavy metadata
- fast app development
Less ideal if:
- you need strict relational querying or complex joins
- you want strong SQL-style tenant/access policies
Best for high-performance / self-hosted: Qdrant
Use this if:
- You want a dedicated vector DB
- You need excellent metadata filtering
- You want good tenant segmentation patterns
- You’re okay running a specialized database
Why it’s strong:
- Fast vector search
- Strong payload filtering
- Can support per-tenant collections or payload-based filtering
- Good balance of simplicity and performance
Best if you need enterprise search features: Weaviate
Use this if:
- You want a rich vector search platform
- You want hybrid search and schema-based metadata
- You like built-in abstractions
Good for:
- semantically rich apps
- hybrid retrieval
Tradeoff:
- More system-specific concepts to learn
My recommendation
Choose PostgreSQL + pgvector if:
- you want metadata filters + tenant access with the fewest moving parts
- you need strong transactional consistency
- your scale is not extreme
Choose Qdrant or Pinecone if:
- embeddings are the core workload
- you expect larger vector scale or higher QPS
- you want a purpose-built vector database
Tenant access patterns to consider
For multi-tenant access, common patterns are:
-
Shared index/table with
tenant_idfilter- simplest
- works well with pgvector, Qdrant, Pinecone, Weaviate
-
Separate namespace/collection/index per tenant
- stronger isolation
- more operational overhead
- good for enterprise or regulated workloads
-
Database-level isolation
- strongest separation
- usually only for high-security cases
If tenants should never see each other’s data, I’d strongly consider:
- Postgres with RLS, or
- separate collections/namespaces in a vector DB plus app-enforced auth
Quick decision rule
- Need simplicity and SQL filters? → Postgres + pgvector
- Need managed vector scaling? → Pinecone
- Need strong filtering and good performance self-hosted? → Qdrant
- Already on MongoDB and metadata is document-shaped? → MongoDB Atlas Vector Search
If you want, I can also give you:
- a decision table comparing these options,
- a reference architecture for multi-tenant embeddings, or
- a Postgres schema example with
pgvectorand tenant RLS.