Prompt
best RAG architecture for internal documents
Latest observation
For internal documents, the best RAG architecture is usually a hybrid, permission-aware, multi-stage retrieval pipeline rather than a simple “embed chunks and search” setup.
Recommended architecture
1) Document ingestion layer
- Connect to sources: Google Drive, SharePoint, Confluence, Notion, Slack, PDF folders, wikis, ticketing systems
- Normalize content into a common schema:
doc_id,title,body,source,author,timestamp,department,acl/permissions,url
- Extract structure:
- headings, tables, lists, code blocks, metadata
2) Chunking strategy
Use structure-aware chunking, not fixed-size only.
- Chunk by section/heading first
- Keep chunks around 300–800 tokens depending on document type
- Add overlap only when needed
- Preserve parent context:
- section title
- document title
- breadcrumb path
For complex docs, use hierarchical chunks:
- small chunk for retrieval
- larger parent chunk for answer generation
3) Indexing layer: hybrid retrieval
Use both:
- Dense vectors for semantic search
- Sparse/BM25 for exact term matching, acronyms, IDs, names, error codes
This is especially important for internal docs because users often search for:
- product names
- policy numbers
- system IDs
- acronyms
- error messages
4) Metadata + ACL filtering
This is critical for internal use.
- Store permissions at document and chunk level
- Filter retrieval by:
- user role
- team
- region
- document confidentiality
- Enforce security before generation, not after
5) Retrieval pipeline
Best practice is a multi-stage retrieval flow:
-
Query understanding
- rewrite query
- expand acronyms
- detect intent and entities
-
Hybrid candidate retrieval
- BM25 + vector search
- retrieve top 50–200 candidates
-
Reranking
- use a cross-encoder or LLM reranker
- rerank top candidates to top 5–10
-
Context assembly
- deduplicate
- prefer diverse sources
- include parent context and citations
6) Generation layer
- Use a strong LLM with:
- grounded answering
- citation support
- refusal when evidence is insufficient
- Ask it to:
- answer only from retrieved context
- cite sources
- mention uncertainty if needed
7) Feedback and evaluation
Track:
- retrieval precision/recall
- answer faithfulness
- citation accuracy
- user satisfaction
- “no answer” correctness
Use logged queries to improve:
- chunking
- synonyms
- query rewriting
- reranking
Best-practice architecture pattern
If you want the shortest recommendation:
Hybrid RAG + reranker + ACL-aware retrieval + hierarchical chunks
This is the most reliable architecture for internal documents.
Why this is better than simple vector RAG
Simple vector-only RAG often fails on internal docs because:
- acronyms and exact terms matter
- docs are messy and unstructured
- permissions must be enforced
- duplicate/near-duplicate docs are common
- query intent is often narrow and factual
Hybrid retrieval plus reranking solves most of that.
Common stack choices
- Embeddings: OpenAI text-embedding, Cohere, Voyage, BGE
- Vector DB: Pinecone, Weaviate, Milvus, pgvector
- Sparse search: Elasticsearch/OpenSearch
- Reranker: Cohere rerank, BGE reranker, cross-encoder model
- Orchestration: LangChain, LlamaIndex, custom service
- Auth/ACL: app-level permission filters integrated into retriever
If your docs are very large or hierarchical
Use Graph RAG or hierarchical RAG when:
- documents reference other documents heavily
- policies depend on many linked pages
- you need traceability across departments
But for most internal knowledge bases, hybrid hierarchical RAG is the best default.
Practical default recommendation
If I were designing this for an enterprise internal knowledge assistant:
- Ingest all docs into a normalized store
- Parse structure and metadata
- Chunk by sections with parent links
- Index in:
- vector store
- BM25 search index
- Apply ACL filters at retrieval time
- Retrieve top 100 via hybrid search
- Rerank to top 5–10
- Generate with citations
- Monitor and iterate using real queries
If you want, I can also give you:
- a reference architecture diagram,
- a production-ready tech stack, or
- a step-by-step implementation plan for your use case.