Prompt
I'm trying to build a chatbot that answers from SharePoint, Slack, and Salesforce. What architecture should I use if I need permissions-aware retrieval and source citations?
Latest observation
Use a retrieval-augmented generation (RAG) architecture with centralized auth-aware indexing and per-source permission filtering.
Recommended architecture
1) Ingestion layer
Connectors pull content from:
- SharePoint
- Slack
- Salesforce
For each item, store:
- document/message text
- metadata: source, author, timestamp, URL
- ACL / permission metadata: users, groups, roles, channel visibility, record-level sharing, etc.
2) Permission-aware indexing
Store the content in a retrieval backend that supports filtering, such as:
- vector index + metadata filters
- hybrid search (keyword + vector)
Key point: don’t rely on the LLM to enforce permissions.
Enforce permissions at retrieval time by filtering candidates based on the current user’s identity and entitlements.
Typical flow:
- user logs in via SSO/OAuth/SAML
- your app resolves their identity/groups/roles
- query includes those entitlements
- retriever only returns chunks the user is allowed to see
3) Retrieval pipeline
Use a two-stage retrieval pattern:
- Candidate search across all sources
- Permission filter on metadata / ACLs
- Optional reranking for relevance
- Pass only allowed chunks to the model
For higher precision, use:
- chunk-level indexing
- source-specific ACL normalization
- hybrid retrieval
- reranking model
4) Answer generation
The LLM gets:
- retrieved passages
- source metadata
- citation IDs/URLs
Prompt the model to:
- answer only using retrieved evidence
- cite every factual claim
- say “I couldn’t find that in your accessible sources” when needed
5) Citation layer
Keep a mapping from chunk IDs to source references:
- SharePoint page/file URL
- Slack channel + message permalink
- Salesforce object/record URL
Return citations in the response as:
- inline footnotes, or
- endnotes with clickable links
6) Security and compliance
Important design choices:
- row-level / document-level ACL enforcement
- no indexing of secrets without access control
- audit logs for retrieval and response generation
- source-specific tokenization and refresh for permissions changes
- short cache TTLs or permission-aware caches
Best-practice pattern
The safest and most scalable approach is:
AuthN/AuthZ service → permission resolver → retrieval engine with ACL filters → LLM with cited context
This is often called:
- secure RAG
- permission-aware RAG
- ACL-filtered retrieval
If you want a simple reference architecture
User
↓
SSO/Auth (OAuth/SAML)
↓
Permission Resolver (groups/roles/ACLs)
↓
Query Router
↓
Hybrid Retriever (vector + keyword)
↓
ACL Filter / Security Trimmer
↓
Reranker
↓
Context Builder with source IDs
↓
LLM
↓
Answer + citations
Practical implementation tips
- Normalize permissions from each system into a common ACL schema.
- Index chunks, not whole documents, for better citation and retrieval.
- Store both the text and a canonical source URL/permalink.
- Refresh permissions frequently, especially for Slack and Salesforce.
- If a source has complex sharing rules, consider precomputing visibility tokens for fast filtering.
Bottom line
Build a secure RAG system with centralized identity, ACL-aware retrieval, and citation-preserving chunk indexing. That gives you:
- permission-aware answers
- traceable citations
- support for multiple enterprise systems
If you want, I can sketch a reference architecture diagram or recommend a specific stack (e.g., Azure/OpenAI, AWS, GCP, or open-source).