Prompt
I'm building a chatbot over our internal docs; what RAG stack should I use?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For an internal-docs chatbot, I’d recommend a simple, reliable RAG stack first, then add sophistication only if you need it.
My default recommendation
If you want the best balance of speed + quality:
- Ingestion / parsing:
UnstructuredorApache Tikafor PDFs, Office docs, HTML - Chunking: sentence/semantic chunking with metadata preservation
- Embeddings:
text-embedding-3-largeor a strong open-source embedding model likebge-large - Vector store:
Postgres + pgvectorif you want simplicity, orPinecone/Weaviateif you want managed scaling - Retriever: hybrid search if possible (BM25 + vector) with metadata filters
- Reranker: a cross-encoder reranker like
bge-reranker-largeor Cohere Rerank - LLM: a strong chat model with citations enabled
- Orchestration:
LlamaIndexif you want a doc-centric framework,LangChainif you need broader tool orchestration
Recommended stack by phase
1) MVP stack
Good for getting something working quickly:
- Parser: Unstructured
- Store: pgvector
- Embeddings: OpenAI embeddings or bge
- Framework: LlamaIndex
- Reranking: optional at first
- UI: Slack bot / web chat
This is usually enough for a first internal prototype.
2) Production-ready stack
If you need better answer quality and trust:
- Ingestion pipeline with OCR, deduping, ACL metadata, versioning
- Hybrid retrieval: BM25 + vector
- Reranker for top 20–50 results
- Citation-aware prompting
- Access control filtering before retrieval
- Evaluation harness with golden Q&A
- Observability with traces and feedback
3) Enterprise stack
If documents are sensitive or numerous:
- Search layer: Elasticsearch/OpenSearch for BM25 + metadata
- Vector layer: pgvector or a vector DB
- Reranker: dedicated reranker service
- AuthZ: document-level ACLs enforced in retrieval
- Monitoring: Langfuse / Arize / OpenTelemetry
- Eval: RAGAS + custom internal benchmarks
What matters most in practice
The stack matters less than these design choices:
-
Good chunking
- Use semantic boundaries where possible
- Keep chunks small enough for retrieval, large enough for context
- Preserve section titles and source metadata
-
Hybrid retrieval
- Dense vectors alone miss exact terms, acronyms, ticket IDs, policy names
- BM25 helps a lot for internal docs
-
Reranking
- This is often the biggest quality boost after retrieval
- Especially important with noisy docs
-
Metadata and permissions
- Department, doc type, date, owner, ACLs
- Filter before generation, not after
-
Evaluation
- Build a set of 50–200 real internal questions
- Track answer correctness, groundedness, and citation quality
My practical stack pick
If I had to choose one stack today for most teams:
- LlamaIndex
- Unstructured
- Postgres + pgvector
- BM25 via OpenSearch or Postgres full-text
- bge embeddings
- bge-reranker
- Claude / GPT-4.1 / similar strong LLM
- Langfuse for tracing
- RAGAS + internal eval set
That gives you a strong path from prototype to production without overengineering.
When to pick something else
- Use Elasticsearch/OpenSearch if your docs are huge and keyword search is already important
- Use Pinecone/Weaviate if you want managed vector infra and less ops
- Use Haystack if you want a more search/pipeline-oriented framework
- Use Azure AI Search if you’re deep in Microsoft/Azure and want integrated enterprise search
Biggest mistake to avoid
Don’t start with:
- only embeddings + vector search
- huge chunks
- no reranker
- no ACL filtering
- no evaluation set
That usually leads to a chatbot that “looks smart” but answers unreliably.
If you want, I can give you:
- a reference architecture diagram,
- a stack recommendation based on your constraints (cloud, budget, doc types, scale), or
- a starter implementation plan in Python.