Prompt

How do I search across text embeddings with metadata filters?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To search across text embeddings with metadata filters, you typically do a vector similarity search plus a structured metadata filter.

Basic idea

  1. Store each text chunk/document with:

    • its embedding vector
    • metadata like author, date, category, source, tenant_id, etc.
  2. When querying:

    • convert the user query into an embedding
    • search for the nearest vectors
    • apply metadata conditions to restrict the candidate set

Common patterns

1) Filter first, then vector search

Use metadata to narrow the dataset, then compute similarity only on matching records.

Example:

  • “Find similar docs from 2024 in finance”
  • Filter: year = 2024 AND category = 'finance'
  • Then do vector search within that subset

This is common when the metadata is selective and can reduce cost.


2) Vector search first, then filter

Retrieve top-k similar vectors, then discard those that don’t match metadata.

This is simpler but can be less accurate if the filter is restrictive, because good matches may be excluded after retrieval.


3) Combined hybrid query

Many vector databases support both in one query:

  • similarity ranking by embedding distance
  • metadata predicate in the same request

This is usually the best option.


Example query logic

Suppose each record looks like:

{
  "id": "doc_123",
  "text": "How to file expense reports",
  "embedding": [0.12, -0.44, ...],
  "metadata": {
    "department": "finance",
    "created_at": "2024-03-12",
    "language": "en"
  }
}

Query:

  • “expense policy”
  • filter: department = finance AND language = en

The system:

  1. embeds "expense policy"
  2. finds nearest embeddings
  3. keeps only items with matching metadata
  4. returns top results

Example in pseudocode

query_vec = embed("expense policy")

results = vector_search(
    vector=query_vec,
    top_k=10,
    filter={
        "department": "finance",
        "language": "en"
    }
)

Example SQL-like filter

If your system supports SQL-style predicates:

SELECT id, text
FROM documents
WHERE department = 'finance'
  AND created_at >= '2024-01-01'
ORDER BY embedding <-> :query_embedding
LIMIT 10;

Tips for good results

  • Index metadata fields you filter on often.
  • Use high-cardinality filters carefully; they can strongly reduce candidates.
  • Keep embeddings and metadata in the same record or strongly linked.
  • If using chunks, store metadata from the parent document on each chunk.
  • Consider hybrid search (vector + keyword) if exact terms matter.

If you're using a vector DB

Different systems use slightly different syntax:

  • Pinecone: metadata filter in the query
  • Weaviate: where filters
  • Milvus: boolean expressions
  • Qdrant: payload filters
  • pgvector/Postgres: combine vector distance with WHERE
  • Elasticsearch/OpenSearch: kNN with filters or hybrid retrieval

If you want, I can show:

  1. a concrete example in Pinecone, Qdrant, Weaviate, or pgvector, or
  2. a recommended schema for storing embeddings + metadata.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.

How do I search across text embeddings with metadata filters? — Obsurfable Explorer