Prompt
I’m planning a private AI assistant over our company docs, tickets, and CRM. What architecture should I use if I need citations, permissions-aware retrieval, and fresh sync from source systems?
Latest observation
For a private AI assistant over company docs, tickets, and CRM with citations, permissions-aware retrieval, and fresh sync, I’d use a RAG architecture with security trimming and an ingestion pipeline built around source-of-truth systems.
Recommended architecture
1) Source systems
Keep your systems of record as-is:
- Document stores: Google Drive, SharePoint, Confluence, Box, etc.
- Ticketing: Zendesk, Jira, ServiceNow, Linear, etc.
- CRM: Salesforce, HubSpot, Dynamics, etc.
Do not copy these into a “new primary database” for business logic. Treat the AI layer as a retrieval and indexing layer only.
2) Ingestion + sync layer
Build connectors that continuously pull from each source system.
Key requirements:
- Incremental sync using webhooks where available, plus scheduled backfills
- Change detection on content, metadata, and permissions
- Normalization into a common internal schema
- Versioning so you can track what content was indexed at any time
A good pattern is:
- Source connector
- Change queue
- Document processor
- Chunker
- Embedding/indexer
- Metadata/ACL store
3) Common content model
Normalize everything into a shared schema like:
document_idsource_type(doc/ticket/crm)source_systemtitlebodychunksmetadata:- owner
- department
- created_at
- updated_at
- URL
- record type
- tags
- customer/account IDs
acl/visibility:- users
- groups
- roles
- tenant / workspace
citation_fields:- source URL
- section/title
- timestamp
- record version
This is critical for both security trimming and citations.
4) Indexing layer
Use a hybrid retrieval system:
- Keyword index for exact matches, names, ticket IDs, account numbers
- Vector index for semantic search
- Optional reranker for final relevance ordering
Store chunks with:
- chunk text
- embeddings
- document and source metadata
- permission metadata or references to ACL records
For permissions, you can either:
- Filter at query time using ACL metadata, or
- Precompute access scopes and attach them to chunks, or
- Use a post-filtering security layer after retrieval
Best practice is usually:
- Query-time permission filtering
- plus document-level ACL precomputation for speed
- plus source-side permission sync
5) Permission-aware retrieval
This is one of the most important parts.
You need retrieval to return only content the user can access. That means:
At ingestion:
- Pull source ACLs / sharing rules / group memberships
- Map them to your identity provider, e.g. Okta, Azure AD, Google Workspace
- Store permission metadata alongside each document/chunk
At query time:
- Identify the user from SSO/JWT
- Resolve their effective permissions:
- user
- groups
- roles
- account/team/tenant
- Retrieve only chunks/documents matching those permissions
Important:
Do not rely only on the LLM prompt to “not reveal secrets.” The retrieval layer must enforce access control before content reaches the model.
This is often called:
- security trimming
- ACL-filtered retrieval
- permission-aware RAG
6) Freshness / sync strategy
To keep answers up to date:
- Use webhooks for near-real-time updates when supported
- Use delta APIs or modified-since queries
- Run a periodic reconciliation job to catch missed changes
- Re-embed only changed chunks
- Expire or delete indexed content when removed from source
Recommended sync design:
- Real-time events for creates/updates/deletes
- Scheduled consistency sweep every few hours or daily
- Tombstones for deletions
- Versioned indexing so old citations can still be traced if needed
7) Retrieval and answer generation flow
A typical request flow:
- User authenticates via SSO
- System resolves user identity and permissions
- Query is rewritten or expanded if helpful
- Hybrid retrieval runs over allowed content only
- Reranker selects top evidence
- LLM generates answer using retrieved context
- Response includes citations linked to source snippets
8) Citations
If you want useful citations, store source-level references during ingestion.
Each retrieved chunk should carry:
- source system
- source record ID
- source URL
- chunk offset or section heading
- last updated timestamp
- document version/hash
In the answer, cite at the chunk level or sentence level:
[1] Confluence page “Onboarding Policy”, section “Equipment”, updated 2026-09-12[2] Zendesk ticket #48321, comment by Support Ops, 2026-09-14
For best results:
- Keep chunks small enough to cite precisely
- Preserve section headers and hierarchy
- Link citations back to the original system
9) Auditability and observability
For enterprise use, log:
- user identity
- query
- retrieved documents/chunks
- permission checks
- citations returned
- model response
- source version at answer time
This gives you:
- audit trails
- debugging
- compliance evidence
- replayability
10) Security and compliance
You’ll want:
- SSO / OIDC / SAML
- tenant isolation if multi-tenant
- encryption at rest and in transit
- secrets management
- data retention policies
- redaction for PII/PCI/PHI if relevant
- least-privilege service accounts for connectors
- admin controls for indexing scopes
Reference architecture summary
A solid architecture looks like this:
Sources → Connectors / webhooks / delta sync → Normalization + ACL mapping → Chunking + metadata extraction → Hybrid search index (keyword + vector) → Permission filter / security trimming → Reranker → LLM answer generation → Citations + audit logs
Practical implementation choices
If you want a simpler first version
Use:
- One search index with metadata filters
- A vector DB that supports metadata filtering
- A text search engine like Elasticsearch/OpenSearch
- SSO-based ACL resolution
- A document store for source snapshots
This gets you to production quickly.
If you need strong enterprise scale
Use:
- OpenSearch/Elasticsearch for keyword and filters
- Vector DB or vector-enabled search for semantic retrieval
- Separate permissions service
- Event-driven ingestion pipeline
- Queue-based sync processing
- Reranking service
- Audit logging and lineage store
What to avoid
- Letting the LLM browse raw systems directly
- Indexing everything without ACLs
- Relying on prompt instructions for security
- Not storing source references for citations
- Using only embeddings without keyword search
- Ignoring deletes and permission changes
- Syncing only nightly if freshness matters
My recommendation in one sentence
Use a permission-aware hybrid RAG architecture with event-driven sync, normalized content + ACL metadata, keyword + vector retrieval, and source-linked chunk citations.
If you want, I can also sketch:
- a concrete system diagram,
- a database/index schema, or
- a cloud-native implementation on AWS/Azure/GCP.