Prompt
How do I set up hybrid search for RAG with vector and keyword retrieval?
Latest observation
Hybrid search for RAG usually means combining:
- Vector retrieval: finds semantically similar chunks
- Keyword / lexical retrieval: finds exact terms, names, IDs, error codes, etc.
The goal is to get the best of both: semantic match + exact-match precision.
1) Basic architecture
Typical pipeline:
- Chunk documents
- Index each chunk twice
- A dense embedding in a vector index
- A text representation in a keyword index (BM25 / inverted index)
- Query both retrievers
- Merge results
- Optionally rerank
- Send top chunks to the LLM
2) Common implementation options
Option A: One engine that supports both
Use a search system that supports dense + lexical retrieval in one place, such as:
- Elasticsearch / OpenSearch
- Weaviate hybrid search
- Azure AI Search
- Pinecone with sparse+dense or hybrid features
- Qdrant with sparse vectors + dense vectors
- Vespa
This is often easiest in production.
Option B: Two separate retrievers
Use:
- Vector DB for embeddings
- Elasticsearch / BM25 for keyword search
Then combine results in your app.
This is flexible and easy to prototype.
3) Recommended hybrid scoring strategy
Two common ways to combine results:
A. Weighted score fusion
Normalize scores from both retrievers and combine:
final_score = alpha * vector_score + (1 - alpha) * keyword_score
alphaaround0.5is a starting point- Increase
alphaif semantic matching is more important - Decrease
alphaif exact term matching matters more
B. Reciprocal Rank Fusion (RRF)
RRF combines rankings instead of raw scores:
score = Σ 1 / (k + rank)
Why it’s popular:
- Robust across different scoring scales
- Works well when retrievers use very different score distributions
If you’re unsure, start with RRF.
4) What to index
For each chunk, store:
- Chunk text
- Document metadata
- Embedding vector
- Searchable text fields for BM25
Good practice:
- Chunk size: ~200–500 tokens
- Overlap: ~10–20%
- Preserve metadata like title, section, source, page number
If exact matches matter, include:
- Titles
- Headers
- Product names
- Code snippets
- IDs, dates, error codes
5) Query flow
For a user query:
- Create embedding for the query
- Run vector search top
k_vresults - Run keyword/BM25 search top
k_bresults - Merge
- Deduplicate by chunk ID
- Re-rank if needed
- Pass top
nchunks to the LLM
Typical values:
k_v = 10k_b = 10- final context chunks = 4–8
6) Example in pseudocode
query = "How do I reset error code E42 on Model X?"
vector_results = vector_index.search(embed(query), top_k=10)
keyword_results = bm25_index.search(query, top_k=10)
merged = reciprocal_rank_fusion(vector_results, keyword_results)
top_chunks = merged[:5]
context = "\n\n".join(chunk.text for chunk in top_chunks)
answer = llm.generate(f"Use the context below to answer:\n{context}\n\nQuestion: {query}")
7) If using Elasticsearch / OpenSearch
You can do hybrid search by combining:
matchormulti_matchfor lexical retrievalknnfor vector similarityfunction_scoreor a custom rank fusion approach
Example idea:
- lexical query for exact terms
- kNN query for embeddings
- combine the two result sets
Many teams implement RRF at the application layer even if the backend supports both.
8) If using LangChain / LlamaIndex
Both frameworks support hybrid retrieval patterns.
LangChain
You can:
- Use a vector store retriever
- Use a BM25 retriever
- Combine with
EnsembleRetriever
LlamaIndex
Supports:
- Hybrid retrievers
- Rerankers
- Fusion techniques
These are good for fast prototyping.
9) Add a reranker for best quality
Hybrid retrieval gets you candidate passages. A reranker improves final ordering.
Common rerankers:
- Cross-encoder rerankers
- Cohere rerank
- bge-reranker
- Instructor rerankers
Pipeline:
Hybrid retrieval → rerank top 20 → keep top 5
This often gives a big quality boost.
10) Practical tips
- Use hybrid search when queries are mixed: natural language + exact terms
- Use keyword-heavy weighting for:
- error codes
- legal text
- medical terminology
- product docs
- APIs and config values
- Use vector-heavy weighting for:
- conceptual questions
- paraphrases
- vague queries
Also:
- Deduplicate chunks before sending to the LLM
- Filter by metadata when possible
- Log query-to-result performance for tuning
- Evaluate with recall@k and answer quality, not just retrieval scores
11) A simple default setup
If you want a sane starting point:
- Chunk size: 300 tokens
- Overlap: 50 tokens
- Retrieve:
- 10 vector results
- 10 BM25 results
- Merge using RRF
- Rerank top 20
- Send top 5 to the LLM
12) Quick recommendation
If you’re building from scratch:
- Prototype: separate BM25 + vector retriever, fuse with RRF
- Production: use a single search engine that supports both
- Quality: add a reranker after fusion
If you want, I can also give you:
- a Python example using LangChain,
- an Elasticsearch/OpenSearch hybrid search example, or
- a Qdrant/Weaviate hybrid setup.