Prompt
How do I use embeddings for RAG with metadata filters?
Latest observation
To use embeddings for RAG with metadata filters, you usually do this:
- Chunk your documents
- Create embeddings for each chunk
- Store chunks in a vector database along with metadata
- At query time, filter by metadata first or during search
- Retrieve the most semantically similar chunks
- Feed them to the LLM as context
Basic idea
Each chunk might look like this:
{
"id": "doc_123_chunk_4",
"text": "Our refund policy allows returns within 30 days...",
"embedding": [0.12, -0.08, ...],
"metadata": {
"source": "support_docs",
"department": "billing",
"language": "en",
"product": "pro",
"date": "2025-01-10"
}
}
The embeddings are used for semantic similarity, while metadata is used for constraints like:
- only search English docs
- only search a specific product
- only search docs from a date range
- only search documents from a particular team
Typical retrieval flow
1) Embed the user query
query_embedding = embed("How do I cancel my subscription?")
2) Apply metadata filters
Example:
language = "en"product = "pro"
3) Search vector DB with both:
- similarity to query embedding
- metadata conditions
Pseudo-query:
results = vector_db.search(
embedding=query_embedding,
top_k=5,
filter={
"language": "en",
"product": "pro"
}
)
4) Build the prompt
Use the retrieved chunks as context:
Answer the user using only the following context:
[Chunk 1]
...
[Chunk 2]
...
Two common ways to combine filters and embeddings
Option A: Pre-filter then vector search
You first narrow the dataset using metadata, then search embeddings inside that subset.
Best when:
- metadata is very selective
- you have a lot of irrelevant documents
Example:
- find all docs where
department = "legal" - among those, choose the top 5 by embedding similarity
Option B: Vector search with metadata filter
The vector DB handles both together in one query.
Best when:
- your vector DB supports filtered ANN search
- you want simpler code
This is the most common approach.
Example with Pinecone-style filtering
query = "What is the refund window?"
query_embedding = embed(query)
results = index.query(
vector=query_embedding,
top_k=5,
filter={
"source": {"$eq": "support_docs"},
"language": {"$eq": "en"},
"product": {"$in": ["pro", "enterprise"]}
},
include_metadata=True
)
Example with Elasticsearch / hybrid setup
If you're using Elasticsearch or OpenSearch, you might combine:
- a keyword query
- vector similarity
- filters
Example logic:
- filter:
language=en - vector search against embeddings
- optionally boost exact keyword matches
Good metadata to store
Useful metadata fields:
sourcetitleurldoc_typedepartmentproductlanguagecreated_atupdated_attagssecurity_level
Try to keep metadata:
- small
- structured
- useful for filtering
- not too many unique values unless needed
Tips for better RAG with filters
1) Store metadata at chunk level
If a document is split into chunks, keep metadata on each chunk so filters work properly.
2) Don’t overfilter
If your filters are too strict, you may retrieve nothing. Have a fallback:
- relax filters
- search broader corpus
- retrieve more candidates
3) Use hybrid search when needed
Embeddings are great for meaning, but metadata and keyword search help with:
- exact names
- product codes
- error messages
- IDs
4) Make filters part of retrieval policy
For example:
- user asks about a specific product → filter by product
- user locale → filter by language
- internal user → filter by access control
5) Evaluate retrieval quality
Check:
- Did the filtered results still contain the answer?
- Are filters excluding relevant docs?
- Are embeddings ranking the right chunks?
Simple architecture
User query
↓
Embed query
↓
Apply metadata filters
↓
Vector search
↓
Top-k chunks
↓
LLM prompt
↓
Answer
Common pitfall
If you filter by metadata only after retrieving top-k globally, you may lose relevant results because the best matches might have been filtered out.
So for important filters, use filtered retrieval directly rather than post-filtering.
If you want, I can also show:
- a Python example using FAISS + metadata in a side store,
- examples for Pinecone / Weaviate / Qdrant / Chroma, or
- a full RAG pipeline with metadata filters.