Prompt
what should I use for embeddings if I need cross-tenant isolation?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If you need cross-tenant isolation for embeddings, the safest default is:
Use a separate embedding index / vector store per tenant — and, if the data is sensitive or you need strong guarantees, separate encryption keys and even separate physical databases/accounts per tenant.
Recommended patterns
-
Best isolation: one vector DB/index per tenant
- Prevents accidental cross-tenant retrieval.
- Simplifies access control and deletion.
- Good when tenant count is manageable.
-
Stronger isolation: separate storage + separate keys
- Each tenant gets its own:
- embedding index
- metadata store
- encryption key
- This is the usual choice for regulated or high-security workloads.
- Each tenant gets its own:
-
Shared index with tenant_id filters
- Only use if you need scale/cost efficiency and can tolerate weaker isolation.
- Must enforce tenant filtering at every query path.
- Riskier because a bug in filtering can leak data across tenants.
For the embeddings themselves
The embedding model can usually be shared across tenants because embeddings are derived representations, not tenant-specific secrets. What matters for isolation is where you store and query them.
Practical recommendation
- If you’re asking from a security-first standpoint:
Use per-tenant vector indexes + per-tenant metadata isolation + per-tenant keys. - If you’re optimizing for cost and scale:
Shared model, isolated logical partitions, strict authz, and mandatory tenant filters.
Rule of thumb
- High sensitivity / compliance → separate index per tenant
- Moderate sensitivity / many tenants → shared infrastructure, strict logical isolation
- Low sensitivity → shared index with tenant filters may be acceptable
If you want, I can suggest an architecture for:
- Postgres + pgvector
- Pinecone / Weaviate / Qdrant
- multi-tenant SaaS with encryption and key management