Prompt

How do I set up hybrid search with vectors, keyword search, and filters for internal content?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A solid hybrid search setup for internal content usually combines:

  1. Vector search for semantic relevance
  2. Keyword/BM25 search for exact matches, names, IDs, acronyms
  3. Metadata filters for access control and narrowing scope

Here’s a practical way to set it up.


1) Model your content

Store each document/chunk with:

  • id
  • text
  • embedding (vector)
  • metadata fields like:
    • source
    • title
    • department
    • created_at
    • updated_at
    • doc_type
    • tags
    • acl_groups or allowed_users
    • language

For internal search, ACL metadata is essential so results only include content the user can access.


2) Chunk documents before indexing

For long internal docs, split into chunks such as:

  • 200–500 tokens per chunk
  • overlap of 10–20% to preserve context

Each chunk should inherit the parent document’s metadata, especially permissions.

Example chunk record:

{
  "id": "doc123_chunk4",
  "text": "The expense policy requires receipts for purchases over $25...",
  "embedding": [0.12, -0.08, ...],
  "title": "Expense Policy",
  "department": "Finance",
  "doc_type": "policy",
  "acl_groups": ["finance", "all_staff"],
  "created_at": "2024-09-12"
}

3) Create embeddings

Use an embedding model to convert each chunk into vectors.
At query time, embed the user’s query the same way.

You want semantic matching for queries like:

  • “How do I get reimbursed for travel?”
  • “What’s our policy on laptop replacement?”
  • “Where is the PTO form?”

Even if the exact words aren’t in the document, vector search helps find relevant content.


4) Index for keyword + vector search

You have a few common options:

Option A: Search engine with hybrid support

Use a system that supports both:

  • vector similarity
  • full-text/BM25 keyword search
  • metadata filtering

Examples include:

  • Elasticsearch / OpenSearch
  • Azure AI Search
  • Vespa
  • PostgreSQL + pgvector + full-text search
  • Weaviate
  • Pinecone with integrated sparse/dense hybrid features in some setups

Option B: Two indexes

  • one vector index
  • one keyword search index

Then merge/rerank results in your application.

This gives flexibility but is more work.


5) Apply filters before or during retrieval

Filters reduce search space and enforce permissions.

Common filters:

  • user’s group membership
  • department
  • document type
  • date range
  • language
  • source system

Example:

  • User belongs to engineering
  • Search only docs where acl_groups contains engineering or all_staff

This should happen at query time, not after retrieval, to avoid leaking restricted documents.


6) Combine keyword and vector scores

A standard hybrid approach is:

  • run vector search
  • run keyword search
  • merge the candidate lists
  • rerank using a weighted score or learned reranker

Example scoring:

hybrid_score = 0.6 * vector_score + 0.4 * keyword_score

Or use reciprocal rank fusion (RRF), which is simple and effective:

RRF score = Σ 1 / (k + rank)

RRF is useful because it avoids needing to normalize scores from different systems.


7) Rerank the top results

After retrieving, rerank the top 20–100 results with:

  • a cross-encoder reranker
  • an LLM-based relevance judge
  • business rules, if needed

This improves precision, especially for internal content where exact policy wording matters.


8) A typical query flow

  1. User sends query: “travel reimbursement for meals”
  2. System identifies user access groups and filters
  3. Create query embedding
  4. Run:
    • vector search with filters
    • keyword/BM25 search with filters
  5. Merge results
  6. Rerank top candidates
  7. Return top answers/snippets with citations

9) Example architecture

Ingestion pipeline

  • Ingest docs from SharePoint, Confluence, Drive, Slack exports, etc.
  • Clean and normalize text
  • Chunk documents
  • Extract metadata
  • Generate embeddings
  • Index:
    • dense vectors
    • keyword fields
    • metadata fields

Query pipeline

  • Authenticate user
  • Resolve ACL groups
  • Build filter clause
  • Run hybrid retrieval
  • Rerank
  • Return results

10) Example query pseudocode

user_groups = ["engineering", "all_staff"]

filters = {
    "acl_groups": {"$in": user_groups},
    "doc_type": {"$in": ["policy", "faq", "handbook"]}
}

query_embedding = embed("How do I request travel reimbursement?")

vector_results = vector_search(
    embedding=query_embedding,
    filters=filters,
    top_k=50
)

keyword_results = keyword_search(
    query="How do I request travel reimbursement?",
    filters=filters,
    top_k=50
)

combined = merge_results(vector_results, keyword_results, method="rrf")
reranked = rerank("How do I request travel reimbursement?", combined[:20])

return reranked[:10]

11) Best practices for internal content

Access control

  • Enforce document-level permissions at retrieval time
  • Store ACLs in metadata
  • Test with restricted docs to ensure no leakage

Chunking

  • Keep chunks self-contained
  • Include headings and section context
  • Avoid chunks that are too large

Keyword tuning

  • Boost exact title matches
  • Boost recent documents if freshness matters
  • Handle synonyms and acronyms

Vector tuning

  • Use domain-appropriate embeddings if possible
  • Re-embed when content changes materially

Observability

Track:

  • zero-result queries
  • clicked results
  • query-to-result latency
  • access-denied counts
  • top failed intents

12) If you want the simplest reliable setup

A good baseline is:

  • Elastic/OpenSearch or Azure AI Search
  • dense embeddings + BM25
  • metadata filters for ACLs
  • RRF fusion
  • reranker on top 20 results

This is usually enough for strong internal search quality.


If you want, I can also give you:

  1. a Postgres + pgvector implementation,
  2. an Elasticsearch/OpenSearch example,
  3. or a reference architecture for RAG over internal docs.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.