Prompt

How do I combine keyword search and vector search for internal docs?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A good pattern is hybrid search: use keyword search for exact terms and vector search for semantic matches, then merge/rerank the results.

Why combine them

  • Keyword search is best for:
    • exact names, IDs, error codes
    • rare terms, acronyms
    • filters like title, tags, author, date
  • Vector search is best for:
    • natural-language questions
    • paraphrases and concept matches
    • cases where the user doesn’t know the exact wording

For internal docs, hybrid search usually gives the best relevance.

Common architectures

1) Parallel retrieval + fusion

Run both searches independently:

  • top k results from BM25/keyword
  • top k results from vector similarity

Then combine using a fusion method:

  • Reciprocal Rank Fusion (RRF): simple, strong default
  • weighted score blending: final = a * keyword_score + b * vector_score
  • reranker after merging: cross-encoder or LLM reranker

Best when: you want robustness and easy implementation.

2) Keyword first, vector fallback

  1. Search with keywords.
  2. If results are weak or few, broaden with vector search.

Best when: exact lookup is critical and you want precision first.

3) Vector first, keyword filters

Use vector search to retrieve candidates, then apply keyword constraints or boosting:

  • mandatory terms
  • must-match titles/tags
  • boosting exact phrase matches

Best when: users ask in natural language but certain terms must be preserved.

Practical recommended setup

For internal docs, I’d start with:

  1. Chunk docs into passages of ~200–500 tokens
  2. Index each chunk in:
    • a text index for BM25/keyword
    • an embedding index for vector search
  3. At query time:
    • run both searches in parallel
    • merge using RRF
    • optionally rerank top 20–50 with a cross-encoder
  4. Return:
    • passage + doc title + section + link
    • highlight matched keywords when available

Simple ranking strategy

A solid default is:

  • retrieve 20 from keyword search
  • retrieve 20 from vector search
  • compute RRF score:

[ score(d) = \sum_{s \in {keyword, vector}} \frac{1}{k + rank_s(d)} ]

where k is often 60.

This avoids needing to normalize BM25 and vector scores, which are on different scales.

When to use weights instead

If you can calibrate scores, use weighted blending:

  • 0.6 * vector + 0.4 * keyword for conversational queries
  • 0.7 * keyword + 0.3 * vector for exact/internal terminology

But score normalization is tricky, so RRF is usually easier and more reliable.

Add metadata filters

For internal docs, filters are very useful:

  • department
  • doc type
  • access permissions
  • date range
  • project/team
  • source system

Apply filters before ranking or during retrieval to reduce noise and enforce permissions.

Reranking helps a lot

After hybrid retrieval, use a reranker to reorder the top candidates based on the full query and document chunk text.

This is especially helpful for:

  • long internal docs
  • similar-looking policy pages
  • queries with subtle intent

Example flow

User query: “What’s our policy on parental leave for contractors?”

  • keyword search finds docs with “parental leave”, “contractor”, “policy”
  • vector search finds relevant HR policy pages even if they say “non-employee leave benefits”
  • fusion merges both sets
  • reranker promotes the most directly relevant section

Implementation options

Depending on your stack:

  • Elasticsearch/OpenSearch: BM25 + kNN/hybrid query
  • Postgres + pgvector: text search + vector similarity
  • Vespa: strong native hybrid ranking
  • Pinecone/Weaviate/Qdrant: vector search plus keyword/filtering depending on product
  • LlamaIndex/LangChain: orchestrate hybrid retrieval, then rerank

Best practices

  • Chunk carefully; don’t index huge documents as one blob
  • Store metadata with each chunk
  • Deduplicate near-identical chunks
  • Track clickthrough/feedback to tune weights
  • Evaluate on real internal queries
  • Ensure permissions are enforced before returning results

If you want, I can also give you:

  1. a reference architecture,
  2. a Python example, or
  3. an example using Elasticsearch/OpenSearch, Postgres+pgvector, or LangChain/LlamaIndex.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.