Prompt
How do I improve retrieval quality in an enterprise chatbot?
Latest observation
Improving retrieval quality in an enterprise chatbot usually comes down to four things: better content, better chunking, better search/ranking, and better evaluation.
1) Start with the data
Retrieval can only be as good as the content you index.
- Remove duplicates and stale versions of documents.
- Fix OCR and formatting issues in PDFs, scans, and slide decks.
- Add metadata like department, date, document type, product, region, and access level.
- Keep source freshness visible so newer policies/docs rank ahead of old ones.
- Control permissions at retrieval time so users only see allowed content.
2) Chunk documents well
Bad chunking is a common cause of poor retrieval.
- Chunk by meaning, not just fixed token length.
- Preserve headings, titles, and section hierarchy.
- Use overlap so important context isn’t split across chunks.
- Avoid chunks that are too large to be precise or too small to be meaningful.
- For structured content, consider section-aware or semantic chunking.
A good chunk often includes:
- title
- section heading
- body text
- source metadata
3) Use hybrid retrieval
Pure vector search often misses exact terms; pure keyword search misses semantic matches.
A strong setup is usually:
- Keyword/BM25 retrieval for exact matches, IDs, acronyms, part numbers
- Vector retrieval for semantic similarity
- Hybrid fusion to combine both result sets
This is especially useful in enterprise settings where users ask:
- “What’s our parental leave policy?”
- “What does SSO stand for in our internal docs?”
- “Show me the SOC 2 exception process”
4) Re-rank results
First-stage retrieval is usually broad; re-ranking improves precision.
- Use a cross-encoder or LLM-based reranker on the top 20–100 candidates.
- Rank by query-document relevance, not just embedding similarity.
- Consider boosting:
- exact title matches
- recent docs
- authoritative sources
- documents from the right department
This often gives a big quality jump.
5) Improve query understanding
Users don’t always phrase questions the same way the docs do.
- Query rewrite/expansion: expand acronyms, normalize product names, resolve aliases.
- Intent detection: distinguish policy lookup vs troubleshooting vs how-to.
- Entity extraction: product names, teams, dates, locations.
- Conversation-aware retrieval: use chat history carefully to resolve pronouns and context.
Example:
- User: “What’s the process?”
- Chat history: “for requesting access to production logs”
- Retrieval should expand to that full intent.
6) Add document routing or filtering
Don’t search everything for every query if you can avoid it.
- Route queries to the most relevant corpus:
- HR
- IT
- Legal
- Engineering
- Sales enablement
- Filter by metadata such as:
- geography
- language
- document type
- recency
- audience
This reduces noise and improves precision.
7) Tune for enterprise-specific language
Enterprise content often includes jargon, acronyms, and internal names.
- Build an acronym dictionary
- Maintain synonym mappings
- Use domain-specific embeddings or fine-tuned models if needed
- Add canonical names for products, teams, and systems
8) Evaluate retrieval directly
Don’t rely only on end-to-end chatbot answers.
Measure retrieval with:
- Recall@K: does the right document appear in top K?
- MRR / nDCG: how high does the correct result rank?
- Precision@K: how many top results are relevant?
- Answer coverage: is the needed evidence retrieved at all?
Create a test set from real user queries:
- question
- expected source docs
- relevant passage
- acceptable synonyms
Then track quality over time.
9) Use feedback loops
- Let users mark results as helpful/unhelpful
- Log failed queries and “no answer” cases
- Review which sources get clicked or ignored
- Use these signals to improve chunking, metadata, and ranking
10) Handle permissions and freshness carefully
In enterprise settings, the “best” result must also be:
- accessible
- current
- policy-compliant
A highly relevant document is useless if the user cannot access it, or if it’s outdated.
A practical baseline architecture
A strong enterprise retrieval pipeline often looks like this:
-
Query understanding
- rewrite/expand query
- apply filters from context
-
Hybrid retrieval
- BM25 + vector search
-
Reranking
- score top candidates with a reranker
-
Context assembly
- deduplicate chunks
- diversify sources
- keep citations
-
Generation
- answer only from retrieved evidence
- cite sources
- say “I don’t know” if evidence is weak
Fastest wins
If you want the biggest improvements quickly, start here:
- Add metadata and permissions filtering
- Use hybrid search
- Improve chunking
- Add reranking
- Create an evaluation set from real queries
If you want, I can also give you:
- a retrieval quality checklist
- a reference architecture
- or a tuning plan for your specific stack (e.g., Elasticsearch, OpenSearch, Pinecone, Weaviate, Azure AI Search, etc.).