Prompt

How do I set up hybrid search with keyword and vector matching?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

Hybrid search combines keyword matching and vector similarity so you can find results that are both textually relevant and semantically relevant.

Basic setup

1) Prepare your data

For each document, store:

  • Text fields for keyword search
  • Embeddings/vectors for semantic search
  • Optional metadata like category, timestamp, permissions

Example document:

{
  "id": "doc1",
  "title": "How to reset your password",
  "body": "To reset your password, go to settings...",
  "embedding": [0.12, -0.04, ...]
}

2) Index both representations

You need:

  • A full-text index for keyword matching
  • A vector index for nearest-neighbor search

Depending on your stack, this may be:

  • Elasticsearch / OpenSearch
  • PostgreSQL + pgvector + FTS
  • Pinecone / Weaviate / Milvus / FAISS plus a keyword engine
  • A custom search layer combining both

3) Query with both methods

At query time:

  1. Build an embedding for the user query
  2. Run a keyword search
  3. Run a vector search
  4. Merge or rerank the results

Common fusion strategies

A) Weighted score fusion

Combine keyword and vector scores:

final_score = α * keyword_score + (1 - α) * vector_score
  • α closer to 1 favors exact text matches
  • α closer to 0 favors semantic similarity

Good when both scores are normalized comparably.

B) Reciprocal Rank Fusion (RRF)

A robust approach when score scales differ:

final_score = Σ 1 / (k + rank_i)
  • Merge ranked lists from keyword and vector search
  • Works well in practice because it uses ranks, not raw scores

C) Two-stage retrieval + reranking

  1. Retrieve top N from keyword search
  2. Retrieve top N from vector search
  3. Union results
  4. Rerank with a cross-encoder / LLM reranker

Best quality, slightly more latency.


Example architecture

Option 1: Elasticsearch/OpenSearch

  • Use BM25 for keyword search
  • Use kNN / dense vector field for vector search
  • Use hybrid query or rerank in application logic

Option 2: PostgreSQL

  • Use tsvector + GIN index for keyword search
  • Use pgvector for embeddings
  • Merge results in your app

Option 3: Vector database + separate keyword engine

  • Use a vector DB for semantic search
  • Use Elasticsearch for keyword search
  • Fuse in an API layer

Simple implementation pattern

Step 1: Store embeddings

Generate embeddings with your model of choice and save them with the document.

Step 2: Query both indexes

  • Keyword query: "reset password"
  • Vector query: embedding of "I can’t log into my account"

Step 3: Normalize scores

Bring both score types to a common scale, or use ranking-based fusion.

Step 4: Merge results

Choose top-k by final score.


Example pseudocode

keyword_results = keyword_search(query_text, top_k=20)
vector_results = vector_search(query_embedding, top_k=20)

combined = fuse(keyword_results, vector_results)  # e.g. RRF or weighted sum
ranked = sorted(combined, key=lambda x: x.score, reverse=True)
return ranked[:10]

Tips

  • Use keyword search for exact terms, names, IDs, product codes
  • Use vector search for paraphrases and conceptual matches
  • Add metadata filters before fusion when possible
  • Consider reranking for best relevance
  • Tune the balance between keyword and vector with real queries

If you want, I can show you a concrete hybrid search example for a specific stack like Elasticsearch, Postgres + pgvector, or Python with FAISS.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.