Prompt
vector search with keyword hybrid
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
A hybrid search system combines:
- Keyword search: matches exact terms, stems, filters, phrases
- Vector search: finds semantically similar documents using embeddings
This usually gives better results than either alone.
Typical architecture
-
Index documents twice
- Lexical index for keywords (e.g. BM25 / inverted index)
- Vector index for embeddings (e.g. FAISS, Milvus, Pinecone, pgvector, Elasticsearch kNN)
-
At query time
- Convert query to embedding
- Run keyword retrieval and vector retrieval in parallel
- Combine/rerank results
Common fusion strategies
1) Weighted score fusion
Normalize scores from both systems and combine:
[ score = \alpha \cdot score_{vector} + (1 - \alpha) \cdot score_{keyword} ]
Good when you can calibrate scores.
2) Reciprocal Rank Fusion (RRF)
Very common because it’s simple and robust:
[ RRF(d) = \sum_{i \in systems} \frac{1}{k + rank_i(d)} ]
Usually works well even if the score scales differ.
3) Two-stage retrieval + rerank
- Retrieve top N from keyword and top N from vector search
- Merge candidates
- Rerank with a cross-encoder / LLM / learning-to-rank model
This often gives the best quality.
When hybrid helps
- Queries with exact entities: names, IDs, SKUs, error codes
- Queries with vague intent: “best way to reduce latency”
- Long-tail vocabulary or synonyms
- Sparse metadata and semantic relevance together
Example
Query: apple battery replacement cost
- Keyword search catches docs with “battery replacement” and “cost”
- Vector search catches docs about “repair pricing” or “service fees” even if wording differs
- Hybrid combines both, improving recall and precision
Implementation options
Elasticsearch / OpenSearch
- BM25 + dense vector kNN
- Use function score, RRF, or reranking
PostgreSQL
tsvector/ full-text search +pgvector- Combine with SQL ranking logic
Dedicated vector DB + search engine
- Example: Elasticsearch for keyword, Pinecone/Milvus for vector
- Merge results in application layer
Practical tips
- Use filters first (tenant, permissions, date)
- Normalize scores before weighted fusion
- Evaluate with Recall@K, nDCG@K, MRR
- Tune retrieval sizes separately for keyword and vector results
- Cache embeddings for frequent queries
If you want, I can also provide:
- a Python example of hybrid search,
- an Elasticsearch/OpenSearch query, or
- a pgvector + full-text search example.