Prompt

How do I use embeddings for RAG with metadata filters?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To use embeddings for RAG with metadata filters, you usually do this:

  1. Chunk your documents
  2. Create embeddings for each chunk
  3. Store chunks in a vector database along with metadata
  4. At query time, filter by metadata first or during search
  5. Retrieve the most semantically similar chunks
  6. Feed them to the LLM as context

Basic idea

Each chunk might look like this:

{
  "id": "doc_123_chunk_4",
  "text": "Our refund policy allows returns within 30 days...",
  "embedding": [0.12, -0.08, ...],
  "metadata": {
    "source": "support_docs",
    "department": "billing",
    "language": "en",
    "product": "pro",
    "date": "2025-01-10"
  }
}

The embeddings are used for semantic similarity, while metadata is used for constraints like:

  • only search English docs
  • only search a specific product
  • only search docs from a date range
  • only search documents from a particular team

Typical retrieval flow

1) Embed the user query

query_embedding = embed("How do I cancel my subscription?")

2) Apply metadata filters

Example:

  • language = "en"
  • product = "pro"

3) Search vector DB with both:

  • similarity to query embedding
  • metadata conditions

Pseudo-query:

results = vector_db.search(
    embedding=query_embedding,
    top_k=5,
    filter={
        "language": "en",
        "product": "pro"
    }
)

4) Build the prompt

Use the retrieved chunks as context:

Answer the user using only the following context:

[Chunk 1]
...

[Chunk 2]
...

Two common ways to combine filters and embeddings

Option A: Pre-filter then vector search

You first narrow the dataset using metadata, then search embeddings inside that subset.

Best when:

  • metadata is very selective
  • you have a lot of irrelevant documents

Example:

  • find all docs where department = "legal"
  • among those, choose the top 5 by embedding similarity

Option B: Vector search with metadata filter

The vector DB handles both together in one query.

Best when:

  • your vector DB supports filtered ANN search
  • you want simpler code

This is the most common approach.


Example with Pinecone-style filtering

query = "What is the refund window?"
query_embedding = embed(query)

results = index.query(
    vector=query_embedding,
    top_k=5,
    filter={
        "source": {"$eq": "support_docs"},
        "language": {"$eq": "en"},
        "product": {"$in": ["pro", "enterprise"]}
    },
    include_metadata=True
)

Example with Elasticsearch / hybrid setup

If you're using Elasticsearch or OpenSearch, you might combine:

  • a keyword query
  • vector similarity
  • filters

Example logic:

  • filter: language=en
  • vector search against embeddings
  • optionally boost exact keyword matches

Good metadata to store

Useful metadata fields:

  • source
  • title
  • url
  • doc_type
  • department
  • product
  • language
  • created_at
  • updated_at
  • tags
  • security_level

Try to keep metadata:

  • small
  • structured
  • useful for filtering
  • not too many unique values unless needed

Tips for better RAG with filters

1) Store metadata at chunk level

If a document is split into chunks, keep metadata on each chunk so filters work properly.

2) Don’t overfilter

If your filters are too strict, you may retrieve nothing. Have a fallback:

  • relax filters
  • search broader corpus
  • retrieve more candidates

3) Use hybrid search when needed

Embeddings are great for meaning, but metadata and keyword search help with:

  • exact names
  • product codes
  • error messages
  • IDs

4) Make filters part of retrieval policy

For example:

  • user asks about a specific product → filter by product
  • user locale → filter by language
  • internal user → filter by access control

5) Evaluate retrieval quality

Check:

  • Did the filtered results still contain the answer?
  • Are filters excluding relevant docs?
  • Are embeddings ranking the right chunks?

Simple architecture

User query
   ↓
Embed query
   ↓
Apply metadata filters
   ↓
Vector search
   ↓
Top-k chunks
   ↓
LLM prompt
   ↓
Answer

Common pitfall

If you filter by metadata only after retrieving top-k globally, you may lose relevant results because the best matches might have been filtered out.

So for important filters, use filtered retrieval directly rather than post-filtering.


If you want, I can also show:

  1. a Python example using FAISS + metadata in a side store,
  2. examples for Pinecone / Weaviate / Qdrant / Chroma, or
  3. a full RAG pipeline with metadata filters.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.