Prompt
Should I use vector search for tenant-aware document retrieval?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
Yes—vector search can be a good fit for tenant-aware document retrieval, but only if you enforce tenant isolation correctly.
When it’s a good choice
Use vector search if you need:
- Semantic search across documents
- Natural language queries like “find contracts about renewal penalties”
- Retrieval over large, unstructured text corpora
- Ranking results by meaning, not just keywords
Tenant-aware requirements
If documents belong to multiple tenants, you should treat tenant identity as a hard filter, not a soft preference.
Recommended pattern
-
Store each chunk/document with metadata like:
tenant_iddocument_idacl / permissionssource_typecreated_at
-
At query time:
- first determine the caller’s allowed
tenant_ids - apply a metadata filter to restrict search to those tenants
- then run vector similarity search within that filtered set
- first determine the caller’s allowed
This avoids cross-tenant leakage.
Important warning
Do not rely on embeddings alone to separate tenants. Similarity search can surface cross-tenant matches if the filter is missing or misconfigured.
Best practice
For most multi-tenant systems:
- Use vector search for retrieval
- Use strict metadata filtering for tenant isolation
- Optionally combine with keyword search or re-ranking
When not to use it
Vector search may be overkill if:
- your retrieval is mostly exact lookup
- documents are small and highly structured
- you only need simple filtering by IDs or tags
Practical architecture
A common setup is:
- Postgres + pgvector for moderate scale
- Dedicated vector DB if scale/latency demands it
- Metadata filters enforced in the retrieval layer
- Row-level security or separate indexes/collections per tenant if you need stronger isolation
If you want, I can also sketch a safe multi-tenant retrieval architecture or compare single shared index vs per-tenant indexes.