Prompt

How do I set up hybrid search for RAG with vector and keyword retrieval?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Hybrid search for RAG usually means combining:

  • Vector retrieval: finds semantically similar chunks
  • Keyword / lexical retrieval: finds exact terms, names, IDs, error codes, etc.

The goal is to get the best of both: semantic match + exact-match precision.

1) Basic architecture

Typical pipeline:

  1. Chunk documents
  2. Index each chunk twice
    • A dense embedding in a vector index
    • A text representation in a keyword index (BM25 / inverted index)
  3. Query both retrievers
  4. Merge results
  5. Optionally rerank
  6. Send top chunks to the LLM

2) Common implementation options

Option A: One engine that supports both

Use a search system that supports dense + lexical retrieval in one place, such as:

  • Elasticsearch / OpenSearch
  • Weaviate hybrid search
  • Azure AI Search
  • Pinecone with sparse+dense or hybrid features
  • Qdrant with sparse vectors + dense vectors
  • Vespa

This is often easiest in production.

Option B: Two separate retrievers

Use:

  • Vector DB for embeddings
  • Elasticsearch / BM25 for keyword search

Then combine results in your app.

This is flexible and easy to prototype.


3) Recommended hybrid scoring strategy

Two common ways to combine results:

A. Weighted score fusion

Normalize scores from both retrievers and combine:

final_score = alpha * vector_score + (1 - alpha) * keyword_score

  • alpha around 0.5 is a starting point
  • Increase alpha if semantic matching is more important
  • Decrease alpha if exact term matching matters more

B. Reciprocal Rank Fusion (RRF)

RRF combines rankings instead of raw scores:

score = Σ 1 / (k + rank)

Why it’s popular:

  • Robust across different scoring scales
  • Works well when retrievers use very different score distributions

If you’re unsure, start with RRF.


4) What to index

For each chunk, store:

  • Chunk text
  • Document metadata
  • Embedding vector
  • Searchable text fields for BM25

Good practice:

  • Chunk size: ~200–500 tokens
  • Overlap: ~10–20%
  • Preserve metadata like title, section, source, page number

If exact matches matter, include:

  • Titles
  • Headers
  • Product names
  • Code snippets
  • IDs, dates, error codes

5) Query flow

For a user query:

  1. Create embedding for the query
  2. Run vector search top k_v results
  3. Run keyword/BM25 search top k_b results
  4. Merge
  5. Deduplicate by chunk ID
  6. Re-rank if needed
  7. Pass top n chunks to the LLM

Typical values:

  • k_v = 10
  • k_b = 10
  • final context chunks = 4–8

6) Example in pseudocode

query = "How do I reset error code E42 on Model X?"

vector_results = vector_index.search(embed(query), top_k=10)
keyword_results = bm25_index.search(query, top_k=10)

merged = reciprocal_rank_fusion(vector_results, keyword_results)

top_chunks = merged[:5]
context = "\n\n".join(chunk.text for chunk in top_chunks)

answer = llm.generate(f"Use the context below to answer:\n{context}\n\nQuestion: {query}")

7) If using Elasticsearch / OpenSearch

You can do hybrid search by combining:

  • match or multi_match for lexical retrieval
  • knn for vector similarity
  • function_score or a custom rank fusion approach

Example idea:

  • lexical query for exact terms
  • kNN query for embeddings
  • combine the two result sets

Many teams implement RRF at the application layer even if the backend supports both.


8) If using LangChain / LlamaIndex

Both frameworks support hybrid retrieval patterns.

LangChain

You can:

  • Use a vector store retriever
  • Use a BM25 retriever
  • Combine with EnsembleRetriever

LlamaIndex

Supports:

  • Hybrid retrievers
  • Rerankers
  • Fusion techniques

These are good for fast prototyping.


9) Add a reranker for best quality

Hybrid retrieval gets you candidate passages. A reranker improves final ordering.

Common rerankers:

  • Cross-encoder rerankers
  • Cohere rerank
  • bge-reranker
  • Instructor rerankers

Pipeline:

Hybrid retrieval → rerank top 20 → keep top 5

This often gives a big quality boost.


10) Practical tips

  • Use hybrid search when queries are mixed: natural language + exact terms
  • Use keyword-heavy weighting for:
    • error codes
    • legal text
    • medical terminology
    • product docs
    • APIs and config values
  • Use vector-heavy weighting for:
    • conceptual questions
    • paraphrases
    • vague queries

Also:

  • Deduplicate chunks before sending to the LLM
  • Filter by metadata when possible
  • Log query-to-result performance for tuning
  • Evaluate with recall@k and answer quality, not just retrieval scores

11) A simple default setup

If you want a sane starting point:

  • Chunk size: 300 tokens
  • Overlap: 50 tokens
  • Retrieve:
    • 10 vector results
    • 10 BM25 results
  • Merge using RRF
  • Rerank top 20
  • Send top 5 to the LLM

12) Quick recommendation

If you’re building from scratch:

  • Prototype: separate BM25 + vector retriever, fuse with RRF
  • Production: use a single search engine that supports both
  • Quality: add a reranker after fusion

If you want, I can also give you:

  1. a Python example using LangChain,
  2. an Elasticsearch/OpenSearch hybrid search example, or
  3. a Qdrant/Weaviate hybrid setup.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.