Prompt

How do I build semantic search with embeddings and metadata filters?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

To build semantic search with embeddings + metadata filters, think of it as a two-stage system:

  1. Find semantically similar items using vector embeddings
  2. Restrict or rerank results using structured metadata filters

1) Core idea

Each document/item gets:

  • a text embedding: a vector representation of its meaning
  • metadata: structured fields like:
    • category
    • author
    • created_at
    • language
    • tenant_id
    • price
    • tags

At query time:

  • embed the user query
  • compare the query vector to stored vectors
  • apply metadata filters such as:
    • category = "support"
    • created_at >= 2024-01-01
    • tenant_id = 123

2) Typical architecture

Ingestion

  1. Collect documents
  2. Split long docs into chunks
  3. Generate embeddings for each chunk
  4. Store:
    • chunk text
    • embedding vector
    • metadata
    • document ID / chunk ID

Query

  1. User enters a search query
  2. Embed the query
  3. Search vector index for nearest neighbors
  4. Apply metadata filters
  5. Return top results

3) Data model

A record might look like:

{
  "id": "chunk_123",
  "text": "How to reset your password...",
  "embedding": [0.12, -0.44, 0.98, ...],
  "metadata": {
    "doc_id": "doc_45",
    "category": "help_center",
    "language": "en",
    "tenant_id": "acme",
    "created_at": "2025-01-10"
  }
}

4) How filtering works

There are two common approaches:

A. Pre-filter then vector search

Apply metadata constraints first, then run vector similarity only on matching items.

Pros

  • precise filtering
  • good for strong constraints like tenant_id, language

Cons

  • may reduce recall if too restrictive
  • can be slower if your filter set is large and not well indexed

B. Vector search then post-filter

Find nearest neighbors first, then filter the top candidates.

Pros

  • simpler to implement
  • can work when filters are weak

Cons

  • might miss valid matches that were filtered out after retrieval

Best practice

Use a vector database / search engine that supports hybrid filtering, so filtering happens efficiently during retrieval.


5) Technologies you can use

Common options:

  • Postgres + pgvector
  • Pinecone
  • Weaviate
  • Qdrant
  • Milvus
  • Elasticsearch / OpenSearch with vector support
  • FAISS plus your own metadata store

If you need metadata filters, choose a system that supports them natively.


6) Example with pgvector

Table schema

CREATE TABLE documents (
  id bigserial PRIMARY KEY,
  content text,
  embedding vector(1536),
  category text,
  language text,
  tenant_id text,
  created_at timestamp
);

Query with filter + similarity

SELECT id, content
FROM documents
WHERE category = 'help_center'
  AND language = 'en'
ORDER BY embedding <-> '[0.1, 0.2, ...]'::vector
LIMIT 10;

Here:

  • embedding <-> query_vector = distance operator
  • WHERE clause applies metadata filters

For performance, add indexes on metadata columns and a vector index on embeddings.


7) Example with a vector DB filter

In Pinecone-style pseudocode:

results = index.query(
    vector=query_embedding,
    top_k=10,
    filter={
        "category": {"$eq": "help_center"},
        "language": {"$eq": "en"},
        "tenant_id": {"$eq": "acme"}
    }
)

8) Hybrid search: semantic + keyword

Semantic search is often better when combined with keyword matching.

A strong pattern is:

  • lexical score from BM25 / full-text search
  • semantic score from embeddings
  • metadata filter for narrowing results

This helps when:

  • exact product names matter
  • rare terms matter
  • users search with short queries

9) Chunking matters

For long documents, split into chunks like:

  • 200–500 tokens per chunk
  • slight overlap, e.g. 10–20%

Store chunk metadata:

  • document ID
  • section title
  • page number
  • chunk order

This improves retrieval precision.


10) Practical design tips

Use metadata for hard constraints

Examples:

  • tenant isolation
  • language
  • access permissions
  • date range
  • content type

Use embeddings for soft meaning

Examples:

  • “how do I change my password?”
  • “reset login credentials”
  • “account access issue”

Keep metadata clean and typed

Use consistent field values:

  • language: "en"
  • not sometimes "english" and sometimes "EN"

Index filter fields

If a field will be filtered often, index it.


11) Common pitfalls

  • Filtering after retrieval only can miss relevant results
  • Poor chunking hurts semantic quality
  • No metadata normalization makes filters unreliable
  • Using embeddings alone for exact-match queries can be weak
  • Not separating tenant data can cause security issues

12) Recommended implementation pattern

If you’re starting from scratch:

Simple stack

  • Postgres
  • pgvector
  • full-text search
  • application-side metadata handling

Larger scale

  • Qdrant / Pinecone / Weaviate
  • embedding pipeline
  • metadata filters in the vector index
  • reranking model for top results

13) Minimal workflow

  1. Embed documents
  2. Store embeddings + metadata
  3. On query, embed query
  4. Search by vector similarity
  5. Apply filters
  6. Return top-k
  7. Optional: rerank results

14) If you want the best quality

Use a 3-step retrieval pipeline:

  1. Filter by permissions / tenant / language
  2. Retrieve top 50–100 with vector similarity
  3. Rerank top candidates with a cross-encoder or LLM

This usually gives better quality than vector search alone.


If you want, I can also show you:

  • a Python example with pgvector
  • a Pinecone/Qdrant example
  • or a reference architecture for production semantic search

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.