Prompt
We need a chatbot that only answers from approved internal sources. What should I use for retrieval, indexing, and access control?
Latest observation
If the chatbot must only answer from approved internal sources, the safest pattern is:
Recommended architecture
Use a RAG system with a controlled document store:
-
Ingestion / indexing
- Pull only approved sources from your internal systems.
- Normalize documents, chunk them, and store metadata:
- source system
- document ID
- owner
- classification
- ACLs / groups allowed
- timestamp / version
- Create embeddings for semantic retrieval.
-
Retrieval
- Use hybrid search:
- keyword / BM25 for exact matches
- vector search for semantic matches
- Apply metadata filtering before results reach the model.
- Retrieve only chunks the user is authorized to see.
- Use hybrid search:
-
Answer generation
- The model should answer only using retrieved context.
- If the context is insufficient, it should say it cannot answer from approved sources.
-
Access control
- Enforce authorization in the retrieval layer, not just in the prompt.
- Use document-level and chunk-level ACLs.
- Filter by user identity, groups, department, clearance, etc.
- Log all retrievals for auditing.
What to use for retrieval
Choose a vector database or search platform that supports filtering and hybrid search:
- Elasticsearch / OpenSearch
Good for hybrid keyword + vector search and metadata filters. - Pinecone Good managed vector search with metadata filtering.
- Weaviate Good vector search and hybrid retrieval.
- Milvus Good for large-scale vector search, usually paired with another search layer for BM25.
- Azure AI Search Strong choice if you’re in Microsoft/Azure and need enterprise auth patterns.
- Postgres + pgvector Fine for smaller setups, but less ideal for sophisticated enterprise search.
If you need strict enterprise controls, OpenSearch/Elasticsearch or Azure AI Search are often the best starting points because they combine search, filters, and governance better than a pure vector store.
What to use for indexing
For indexing, use a pipeline that does:
- document extraction: PDFs, Office docs, wiki pages, tickets
- cleaning and chunking
- deduplication
- metadata enrichment
- embedding generation
- ACL attachment
Tools commonly used:
- LlamaIndex or LangChain for ingestion pipelines
- Haystack for end-to-end retrieval workflows
- Custom ETL jobs if governance is strict
Important:
- keep chunks small enough for precision
- store source citations for every chunk
- version documents so stale content can be removed
- reindex on source changes
What to use for access control
Access control should be enforced at multiple layers:
1. Source approval layer
Only ingest from approved repositories.
2. Identity layer
Integrate with SSO / IdP:
- Okta
- Azure AD / Entra ID
- Google Workspace
- Ping
3. Retrieval authorization layer
Before retrieval, resolve the user’s:
- roles
- groups
- departments
- clearance level
Then filter search results by ACL metadata.
4. Output guardrails
Require citations from retrieved documents only. If no approved evidence exists, return:
- “I don’t have enough approved internal sources to answer that.”
5. Auditing
Log:
- who asked
- what was retrieved
- what sources were used
- what answer was returned
Practical implementation pattern
A strong setup looks like:
- Source systems: SharePoint, Confluence, Drive, internal wiki, ticketing system
- ETL/indexing: custom pipeline or LlamaIndex/Haystack
- Search layer: OpenSearch / Azure AI Search
- Auth: Entra ID / Okta
- LLM: any model behind a strict RAG wrapper
- Policy enforcement: in retrieval + response validation
Best-practice recommendation
If your main requirement is “only approved internal sources,” I’d recommend:
- Azure AI Search + Entra ID if you’re Microsoft-centric
- OpenSearch/Elasticsearch + SSO + ACL filtering for flexible enterprise control
- Add hybrid retrieval, metadata-based authorization, and citation-only generation
Key caution
Do not rely on prompting alone to enforce access control.
The model must never even see unauthorized documents.
If you want, I can also give you:
- a reference architecture diagram, or
- a vendor comparison table for OpenSearch vs Azure AI Search vs Pinecone vs Weaviate.