Prompt
I'm building a chatbot over company docs and need an embeddings strategy
Latest observation
Here’s a solid embeddings strategy for a chatbot over company docs, optimized for reliability, retrieval quality, and maintainability.
1) Start with the right retrieval architecture
For company docs, the best default is usually:
- Chunk documents
- Embed chunks
- Store in a vector database
- Retrieve top-k chunks per query
- Optionally re-rank
- Generate answer grounded in retrieved chunks
This is the standard RAG setup and works well for policies, manuals, knowledge bases, PDFs, wiki pages, tickets, etc.
2) Choose an embedding model based on your constraints
Pick an embedding model with strong semantic retrieval performance, good multilingual support if needed, and reasonable cost.
Good options
- OpenAI text-embedding-3-large: best quality, more expensive
- OpenAI text-embedding-3-small: cheaper, strong baseline
- Cohere Embed v3
- Voyage embeddings
- Sentence-transformers / BGE / E5 if you want self-hosted
Rule of thumb
- If accuracy matters most: use a top-tier hosted model
- If cost/latency matters most: use a smaller hosted model
- If data residency or control matters: self-hosted embeddings
3) Chunking strategy matters as much as the model
Bad chunking can ruin retrieval even with great embeddings.
Recommended chunking
- Chunk by semantic structure when possible:
- headings
- sections
- paragraphs
- bullet lists
- Use token-based chunks with overlap as fallback
Typical settings
- Chunk size: 300–800 tokens
- Overlap: 50–150 tokens
Best practices
- Keep chunks focused on one topic
- Avoid cutting tables or lists in the middle
- Preserve metadata:
- document title
- section heading
- source URL/path
- last updated date
- access permissions
4) Use metadata aggressively
Metadata improves filtering, ranking, and trust.
Store for each chunk:
- doc_id
- chunk_id
- title
- heading path
- source
- department
- created_at / updated_at
- version
- ACL / permissions
- document type
Then use metadata filters like:
- only show HR docs to HR users
- only current policy versions
- only docs from a specific region/business unit
5) Use hybrid retrieval, not just vector search
For enterprise docs, hybrid search is usually better than vector-only.
Combine:
- Dense embeddings for semantic matching
- Sparse retrieval like BM25 for exact terms, acronyms, part numbers, policy names
This helps when users ask:
- “What is the PTO policy?”
- “How do I reset Okta?”
- “What is SOX control 3.2?”
- “Where is the Q4 revenue deck?”
Exact terms and jargon often matter a lot.
6) Add a reranker if quality matters
A reranker can significantly improve final retrieval quality.
Pipeline:
- Retrieve top 20–50 chunks using hybrid search
- Re-rank them with a cross-encoder/reranker
- Send top 3–8 chunks to the LLM
This reduces irrelevant context and improves answer accuracy.
7) Consider multiple embeddings per document type
Different document types benefit from different handling:
- Policy docs: chunk by section, preserve hierarchy
- FAQs: embed question-answer pairs
- Tables: convert rows to text or row-level chunks
- Slides: often need slide-level chunks plus speaker notes if available
- Tickets / chat logs: thread-level chunks or summarized chunks
8) Plan for updates and versioning
Company docs change constantly.
You want:
- incremental re-embedding for changed docs only
- versioned chunks
- soft deletion of obsolete docs
- re-indexing pipeline triggered by doc updates
Also consider a TTL or freshness boost:
- prefer newer docs when content is similar
- down-rank stale documents
9) Evaluate retrieval, not just answer quality
Measure embeddings strategy using retrieval metrics.
Useful metrics
- Recall@k
- MRR
- nDCG
- Answer faithfulness
- Groundedness
- Citation accuracy
Build a test set of real user questions with expected source docs.
What to test
- exact terminology queries
- paraphrases
- ambiguous queries
- multi-hop questions
- acronym-heavy questions
- stale vs current policy distinctions
10) Practical recommended default stack
If you want a strong, simple starting point:
- Chunking: 500 tokens with 100 token overlap, structure-aware
- Embeddings: a strong hosted embedding model
- Index: vector DB with metadata filtering
- Retrieval: hybrid dense + BM25
- Reranking: yes
- Generation: answer only from retrieved context, with citations
- Freshness: prefer latest approved docs
- Access control: enforce at retrieval time
11) Common pitfalls to avoid
- Using huge chunks that mix multiple topics
- Ignoring metadata and document versioning
- Vector-only search in enterprise settings
- Not handling acronyms and exact terms
- Re-embedding everything on every change
- Not testing retrieval on real user questions
- Letting the model answer without citations or grounding
12) If you want the fastest path to production
Do this:
- Ingest docs
- Parse structure
- Chunk by section/paragraph
- Embed chunks
- Store in vector DB with metadata
- Add BM25 hybrid search
- Add reranker
- Build evaluation set
- Iterate on chunking and retrieval
If you want, I can also give you:
- a specific recommended stack for your budget,
- a chunking recipe for PDFs/wiki/HTML, or
- a reference architecture diagram for the whole chatbot pipeline.