Prompt
How do I add vector search to a RAG app without breaking keyword search?
Latest observation
To add vector search to a RAG app without losing keyword search, use a hybrid retrieval setup rather than replacing BM25/lexical search.
Recommended approach: hybrid search
Combine:
- Keyword search for exact matches, names, IDs, rare terms, quotes, code, and fresh terms
- Vector search for semantic similarity, paraphrases, and concept matches
This avoids the common failure mode where pure vector search misses exact strings and pure keyword search misses meaning.
Practical architecture
1) Keep your existing keyword index
If you already have:
- Elasticsearch / OpenSearch
- Postgres full-text search
- Lucene / BM25-based search
keep it as-is. Don’t remove it.
2) Add embeddings for the same chunks
For each document chunk:
- store text
- store metadata
- store embedding vector
You can store vectors in:
- a vector DB like Pinecone, Weaviate, Milvus, Qdrant
- or the same search system if it supports vectors
- or Postgres with pgvector
3) Retrieve from both systems
At query time:
- run keyword search
- run vector search
- merge results
4) Fuse the results
Use one of these common methods:
Simple weighted score fusion
Normalize scores from both retrievers, then combine:
- final_score =
α * keyword_score + (1 - α) * vector_score
Good when your scoring scales are stable.
Reciprocal Rank Fusion (RRF)
This is often the safest default:
- rank documents from each retriever
- combine based on rank positions, not raw scores
RRF is robust when keyword and vector scores are on different scales.
A common formula:
1 / (k + rank)- sum across retrievers
Best practice: use both, then rerank
A strong pattern is:
- Keyword search retrieves top 50
- Vector search retrieves top 50
- Merge to a candidate set
- Rerank top 20–50 with a cross-encoder or LLM reranker
- Send top chunks to the generator
This gives you:
- keyword precision
- vector recall
- better final ordering
How to avoid breaking keyword search
These are the main pitfalls:
Don’t replace lexical retrieval
Pure vector search can fail on:
- product codes
- error codes
- exact legal terms
- names
- numbers
- abbreviations
So always keep keyword retrieval available.
Don’t chunk too aggressively
If chunks are too small:
- keyword matches may lose context
- vector embeddings may become noisy
A good starting point:
- 200–500 tokens per chunk
- overlap of 10–20%
Preserve metadata filters
Apply filters like:
- tenant
- document type
- date
- access control
to both keyword and vector retrieval, otherwise results may diverge.
Use the same chunk IDs across systems
This makes merging and deduplication much easier.
A simple implementation pattern
Data ingestion
For each chunk:
- index text in BM25/keyword index
- generate embedding
- store vector + chunk metadata in vector store
Query flow
- User asks a question
- Query is sent to both retrievers
- Merge hits by doc/chunk ID
- Deduplicate
- Fuse ranks/scores
- Rerank
- Feed top results to the LLM
Example hybrid strategy
If using Elasticsearch + vector DB:
- Elasticsearch handles BM25
- vector DB handles similarity
- application layer merges results
If using a single engine like OpenSearch or Azure AI Search:
- use built-in hybrid retrieval
- simpler ops, less glue code
Which fusion method should you choose?
- RRF: best default, easiest to tune
- Weighted score blend: good if you understand scoring distributions
- Learning-to-rank: best at scale, more engineering effort
If you’re starting out, use: BM25 + vector search + RRF + reranker
Tuning tips
- Start with
top_k = 20–50from each retriever - Deduplicate by chunk ID
- If keyword queries are exact or short, weight keyword more
- If queries are long or descriptive, weight vectors more
- Evaluate on a set of real questions and measure retrieval recall, not just answer quality
Example rule of thumb
A practical routing policy:
- short query with IDs/names/numbers → emphasize keyword search
- long natural-language question → emphasize vector search
- always keep both available
If you want the simplest safe setup
Use this:
- existing keyword search unchanged
- add vector search in parallel
- merge with RRF
- rerank top results
- keep metadata filters identical
That gives you hybrid retrieval without breaking the behavior users already rely on.
If you want, I can show you:
- a reference architecture diagram,
- sample code in Python, or
- how to do this in Elasticsearch/OpenSearch, pgvector, or Pinecone.