Prompt
I'm building a RAG app with permissions-aware retrieval, what tools should I use?
Latest observation
For a permissions-aware RAG app, the best tools are usually a combination of:
- A vector store that supports metadata filtering / ACLs
- An auth system as the source of truth for permissions
- An ingestion pipeline that attaches document-level and chunk-level access metadata
- An app/RAG framework that can apply filters before retrieval
Good tool choices
Vector databases with metadata filters
These are commonly used for permissions-aware retrieval:
- Pinecone — strong metadata filtering, good managed option
- Weaviate — supports filters and hybrid search
- Qdrant — very good payload filtering and flexible ACL-style metadata
- Milvus — supports filtering, good at scale
- OpenSearch / Elasticsearch — good if you want keyword + vector + complex filtering
- Postgres + pgvector — great if your scale is moderate and you want SQL-based ACL filtering
If permissions are important, Qdrant, Pinecone, Weaviate, or Postgres/pgvector are common starting points.
Auth / permission source of truth
Don’t hardcode permissions only in the vector DB. Use a real auth system:
- Auth0
- Okta
- Azure AD / Entra ID
- AWS IAM / Cognito
- Keycloak
- Google Workspace / SSO
- Or your internal RBAC/ABAC service
You want something that can answer:
- who is the user?
- what groups/roles/orgs are they in?
- what documents, tenants, or attributes can they access?
RAG frameworks
These help with orchestration, chunking, retrieval, reranking, etc.:
- LlamaIndex — very good for retrieval pipelines and metadata-aware retrieval
- LangChain — flexible and widely used
- Haystack — solid for search/retrieval-heavy systems
If permissions logic is complex, LlamaIndex + a vector DB with filters is a strong combo.
Rerankers
If you do permission filtering first, then rerank the allowed results:
- Cohere Rerank
- bge-reranker
- Jina reranker
- Cross-encoder rerankers from Hugging Face
Recommended architecture
A common pattern:
- Ingest documents
- Chunk them
- Attach ACL metadata to each chunk
tenant_iduser_idsgroup_idsroleclassificationallowed_departments
- Store chunks in vector DB
- At query time:
- authenticate user
- resolve their permissions/groups
- apply metadata filters in retrieval
- rerank allowed docs
- send top results to the LLM
Important design choice: filter before or after retrieval?
Best practice: filter before retrieval
This is safer because you avoid even retrieving unauthorized content.
Example filter:
tenant_id = current_tenantvisibility IN user_allowed_visibilitiesgroup_ids overlaps user_group_ids
Avoid:
- retrieving everything and filtering only in the prompt
- relying on the LLM to “ignore” unauthorized chunks
That’s not secure.
Suggested stack by use case
Simple, production-friendly
- Postgres + pgvector
- Keycloak/Auth0
- LlamaIndex
- Cohere Rerank or bge-reranker
Good if:
- your data size is moderate
- permissions are mostly SQL/RBAC-style
- you want simplicity
More scalable vector-first
- Qdrant or Pinecone
- Auth0 / Entra / Okta
- LangChain or LlamaIndex
- Reranker
Good if:
- lots of embeddings
- need fast filtered similarity search
- permissions are metadata-driven
Search-heavy hybrid RAG
- OpenSearch/Elasticsearch
- SSO + policy service
- Haystack or LangChain
- Reranker
Good if:
- keyword search matters a lot
- documents are long/legal/enterprise content
- you need hybrid lexical + vector retrieval
Extra security tips
- Enforce permissions at the retrieval layer, not just the UI
- Store permission metadata at the chunk level, not only document level
- Use tenant isolation if multi-tenant
- Log access decisions for audit
- Periodically sync ACL changes
- Consider a policy engine like OPA (Open Policy Agent) if rules are complex
My practical recommendation
If you want a strong default choice:
- LlamaIndex
- Qdrant
- Auth0 or Keycloak
- A reranker like Cohere Rerank or bge-reranker
- Optional: OPA if permissions are complex
If you want, I can also give you:
- a reference architecture diagram,
- a tool comparison table, or
- a sample implementation pattern for permissions-aware retrieval.