Prompt
What should I use for permission-aware retrieval over internal documents?
Latest observation
For permission-aware retrieval over internal documents, the best pattern is usually:
Use a search/RAG stack with authorization filtering, not a standalone vector database alone
Recommended approach
-
Ingest documents into:
- a full-text search index (for exact keyword + hybrid retrieval)
- optionally a vector index (for semantic search)
-
Store document-level permissions as metadata:
- user IDs
- group IDs / roles
- ACLs
- tenant/org ID
- sensitivity labels
-
At query time:
- authenticate the user
- resolve their effective permissions
- apply pre-filtering in retrieval so only authorized docs/chunks are searched or returned
Good implementation options
If you want a pragmatic enterprise setup
- OpenSearch / Elasticsearch with ACL metadata filtering
- Add vector search if needed
- Use hybrid retrieval: keyword + semantic
This is often the safest and easiest way to do permission-aware retrieval at scale.
If you want a managed cloud option
- Azure AI Search
- Google Vertex AI Search / Agent Builder
- AWS OpenSearch Service (with custom auth filtering)
These are good if your documents already live in a cloud ecosystem and you want less infra work.
If you want vector-native retrieval
- Pinecone, Weaviate, Milvus, Qdrant
These can work well, but make sure they support:
- metadata filtering
- per-document ACLs
- tenant isolation
- efficient pre-filtering
Vector DBs alone are usually not enough unless you carefully design the security layer around them.
Key design rule
Don’t rely on the LLM to “respect permissions”
The model should only see already-authorized context.
Permission enforcement must happen before retrieval and before prompt assembly.
Important architecture choices
1. Filter before retrieval, not after
Best practice:
- determine allowed doc IDs / chunks first
- search only within that subset
Avoid:
- retrieving broadly and then trying to redact afterward
Because unauthorized text may already have influenced the model.
2. Store permissions at the chunk level if documents are mixed-sensitivity
If a document contains sections with different access levels, chunk-level ACLs are better than doc-level ACLs.
3. Use ACLs or group-based access, not raw user lists when possible
For scale, it’s better to store:
- groups/roles
- department
- org/tenant
- classification labels
Then expand to effective permissions at query time.
4. Consider document partitioning by tenant or sensitivity
For strong isolation, split indexes by:
- tenant
- business unit
- classification tier
This reduces risk and simplifies filtering.
What I’d recommend in practice
For most internal enterprise search/RAG systems:
- Elasticsearch/OpenSearch + metadata ACL filtering + vector embeddings
- Add a policy service to resolve user permissions
- Use hybrid retrieval
- Retrieve only authorized chunks
- Pass only those chunks to the LLM
If you want the simplest managed path:
- Azure AI Search if you’re in Microsoft 365 / Entra ID land
If you tell me your stack, I can recommend more precisely
For example:
- cloud provider
- whether you need document-level or chunk-level permissions
- expected scale
- whether you need keyword, semantic, or both
- whether your source of truth is Google Drive, SharePoint, S3, Confluence, etc.
If you want, I can also give you a reference architecture or a shortlist by stack (AWS/Azure/GCP/open-source).