Prompt

How do I combine embeddings with keyword search?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A common way to combine embeddings with keyword search is to use a hybrid search pipeline:

1) Run both searches

  • Keyword search: good for exact matches, IDs, names, rare terms, filters.
  • Embedding/vector search: good for semantic similarity and paraphrases.

2) Merge the results

Typical approaches:

A. Weighted score fusion

Give each document:

  • a keyword score
  • a vector similarity score

Then combine them:

[ \text{final_score} = \alpha \cdot \text{vector_score} + (1-\alpha) \cdot \text{keyword_score} ]

Choose α based on your use case:

  • higher α if meaning matters more
  • lower α if exact term matching matters more

B. Reciprocal Rank Fusion (RRF)

If your keyword and vector systems produce ranked lists, combine ranks instead of raw scores:

[ \text{RRF}(d) = \sum_i \frac{1}{k + rank_i(d)} ]

This is popular because it’s robust even when score scales differ.

C. Two-stage retrieval

  1. Use keyword search to get candidates
  2. Re-rank candidates with embeddings

Or the reverse:

  1. Vector search for candidates
  2. Boost exact keyword matches in re-ranking

3) Use metadata filters

Before or during retrieval, filter by:

  • date
  • category
  • language
  • permissions
  • product type

This keeps search relevant and efficient.

4) Practical best practices

  • Normalize scores before blending; keyword and vector scores often live on different scales.
  • Use keyword search for:
    • names
    • codes
    • error messages
    • URLs
    • exact phrases
  • Use embeddings for:
    • paraphrases
    • conceptual matching
    • natural language queries
  • Consider stemming, synonyms, and phrase queries in keyword search.

5) Simple implementation pattern

  1. Build an inverted index for keyword search.
  2. Store embeddings in a vector database or ANN index.
  3. For each query:
    • search both indexes
    • retrieve top N from each
    • merge with RRF or weighted scoring
    • optionally rerank with a cross-encoder/LLM

Example

Query: “reset my account password”

  • Keyword search finds documents with “reset” and “password”
  • Embedding search also finds “forgot credentials” or “change login details”
  • Hybrid search returns both exact and semantically similar results

If you want, I can also show:

  • a Python example
  • an example using Elasticsearch/OpenSearch
  • or a vector DB + BM25 hybrid architecture

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.