Prompt
Need RAG with document permissions
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
Yes — you can absolutely build RAG with document-level permissions, and it’s a common enterprise requirement.
Core idea
At retrieval time, you must filter by what the current user is allowed to see before the model gets the context.
That usually means:
- Ingest documents
- Chunk and embed them
- Store each chunk with permission metadata
- On query, identify the user
- Retrieve only chunks the user is authorized to access
- Send those chunks to the LLM
Common permission models
You can support:
- User-based access: specific users can see specific docs
- Group/team-based access: documents shared with groups
- Role-based access: based on roles like admin, HR, legal
- Attribute-based access: department, region, clearance, tenant, etc.
Recommended architecture
1) Store metadata with each chunk
For each chunk, store fields like:
document_idchunk_idtenant_idallowed_user_idsallowed_group_idsclassificationdepartmentacl_version
2) Enforce access at retrieval
When a user queries:
- resolve their identity
- compute their effective permissions
- use metadata filters in the vector DB / search layer
- only retrieve permitted chunks
3) Keep permissions in a source of truth
Don’t rely only on vector DB metadata as the authority. Have:
- a document permission service or database
- sync permissions into the retrieval index
- periodically reconcile
Retrieval patterns
Pattern A: pre-filter then vector search
Best when your vector DB supports metadata filtering.
- Apply ACL filter
- Search only within allowed chunks
This is usually the best option.
Pattern B: vector search then post-filter
Works if the DB lacks good metadata filtering, but:
- can miss relevant results
- can leak side-channel info
- less secure if not done carefully
Pattern C: hybrid search with security filter
Use keyword + vector search, but always apply ACL constraints.
Important security notes
- Never pass unauthorized chunks to the LLM, even if “hidden” in the prompt
- Apply filters server-side, not in client code
- Log access decisions
- Consider row-level security in your DB
- Re-check permissions if ACLs change
- Encrypt sensitive docs at rest and in transit
Example schema
{
"chunk_id": "c123",
"document_id": "d45",
"text": "Confidential roadmap ...",
"tenant_id": "t1",
"allowed_groups": ["eng", "pm"],
"allowed_users": ["u17"],
"classification": "confidential"
}
Query-time filter:
{
"tenant_id": "t1",
"$or": [
{"allowed_users": "u17"},
{"allowed_groups": {"$in": ["eng"]}}
]
}
Practical implementation options
- Pinecone: metadata filtering
- Weaviate: filtering + tenants
- Milvus: scalar filtering
- OpenSearch / Elasticsearch: strong filtering for hybrid search
- Postgres + pgvector: possible with row-level security
- Azure AI Search / Vertex AI Search: enterprise access controls depending on setup
Best practice
For enterprise RAG, the safest design is:
- authoritative ACL store
- sync ACLs into retrieval metadata
- query-time authorization filter
- audit logging
- tenant isolation
If you want, I can give you:
- a reference architecture diagram
- a Python implementation example
- a Postgres/pgvector design
- a Pinecone/Weaviate example
- a full production-ready ACL + RAG pattern