Prompt
I'm building a chatbot on top of company docs and need a production retrieval layer
Latest observation
A production retrieval layer for company docs should be built like a search system first, and a chatbot feature second. The main goals are: high recall, good ranking, low latency, access control, observability, and easy re-indexing.
Recommended architecture
1) Ingestion pipeline
- Pull from sources: Google Drive, Confluence, SharePoint, Notion, Slack, PDFs, webpages, ticketing systems
- Normalize into a common document schema:
doc_idsourcetitlebodyauthorcreated_atupdated_atacl / permissionstags / department / tenanturl
- Clean text:
- remove boilerplate
- extract tables and headings
- deduplicate near-identical content
- preserve structure where possible
2) Chunking strategy
Use semantic chunking, not fixed-length only.
- Chunk by headings/sections when available
- Target ~300–800 tokens per chunk
- Add overlap only when needed
- Store metadata with every chunk:
- parent document
- section path
- chunk order
- ACL
- source URL
A good rule: chunks should be large enough to carry meaning, small enough to retrieve precisely.
3) Indexing
Use a hybrid index:
- BM25 / keyword index for exact terms, names, codes, acronyms
- Vector index for semantic retrieval
- Optionally a reranker for top-k results
This is usually better than vector-only for enterprise docs.
4) Retrieval flow
A strong production flow:
-
Query understanding
- detect language
- optionally rewrite the query for search
- expand acronyms
- identify filters like department, date, doc type
-
Candidate retrieval
- retrieve top
kfrom BM25 - retrieve top
kfrom vector search - merge results
- retrieve top
-
Filtering
- enforce ACLs before returning candidates
- filter by tenant, department, freshness, doc type if needed
-
Reranking
- use a cross-encoder or LLM-based reranker on top 20–100 candidates
- optimize for relevance to the user’s question, not just similarity
-
Context assembly
- deduplicate overlapping chunks
- keep citations
- limit total context size
- prefer chunks with direct answer evidence
-
Answer generation
- send only retrieved context to the LLM
- require citations
- instruct it to say “I don’t know” if evidence is insufficient
Key production requirements
Access control
This is critical in company docs.
- Enforce permissions at retrieval time, not just in the UI
- Support document-level and chunk-level ACLs if needed
- If permissions differ across sections, store ACL metadata per chunk
Freshness
- Incremental indexing for updated docs
- Support deletions and permission changes
- Track document versioning
- Use freshness scoring if recent docs should rank higher
Observability
Track:
- query volume
- top queries with no good results
- click-through or answer acceptance
- retrieval latency
- reranker latency
- embedding/index freshness
- hallucination reports
- citation coverage
Evaluation
You need offline and online evaluation.
Offline:
- precision@k
- recall@k
- MRR / nDCG
- answer-groundedness
- ACL leakage tests
Online:
- user satisfaction
- answer correctness
- retrieval success rate
- escalation rate
- latency
Create a labeled set of real internal questions and expected supporting passages.
Suggested tech stack
If you want fast and practical
- Postgres + pgvector for smaller corpora
- OpenSearch / Elasticsearch for BM25 + hybrid search
- A reranker model for reordering top results
- Object storage for raw documents and parsed text
If you expect scale
- Elasticsearch/OpenSearch for lexical retrieval
- Dedicated vector DB:
- Pinecone
- Weaviate
- Milvus
- pgvector if scale is moderate
- Separate metadata store for ACLs and audit logs
Practical retrieval recipe
A strong baseline:
- Query both BM25 and vector search
- Combine with reciprocal rank fusion
- Filter by permissions
- Rerank top 50 with a cross-encoder
- Return top 5–10 chunks with citations
This usually outperforms pure vector search.
Common failure modes
- Bad chunking: splitting answers across chunks
- No ACL enforcement: data leakage
- Vector-only search: misses exact terminology
- Too much context: distracts the LLM
- No reranker: irrelevant chunks surface
- Stale index: answers based on outdated docs
- No evaluation set: impossible to know if retrieval is improving
A good production principle
The retriever should answer:
“What passages most likely contain the evidence needed to answer this user’s question, under their permissions, with the latest trusted content?”
If you want, I can also give you:
- a reference architecture diagram,
- a schema design for chunks + ACLs, or
- a concrete implementation using OpenSearch + pgvector + reranking.