Prompt
How do I combine metadata filters with vector similarity search?
Latest observation
To combine metadata filters with vector similarity search, you typically do one of these:
1) Filter first, then search within the filtered set
Use metadata conditions to narrow the candidate documents, then run vector similarity on that subset.
Example pattern:
- Metadata:
category = "legal"andyear >= 2023 - Vector search: find the top 10 most similar embeddings among those documents
This is common when your search system supports a where / filter clause alongside vector search.
Pseudo-query:
{
"vector": [0.12, 0.98, ...],
"top_k": 10,
"filter": {
"category": "legal",
"year": { "$gte": 2023 }
}
}
2) Search first, then filter results
Run vector search over all items, then remove results that don’t match metadata.
This is simpler, but can hurt recall:
- If the top similar items don’t satisfy the filter, they get discarded
- You may end up with fewer than
top_kresults
3) Hybrid/combined retrieval
Some systems combine:
- metadata filtering
- vector similarity
- sometimes keyword/BM25 search
This is often the best approach for production search.
Practical considerations
If your filter is restrictive
Example:
tenant_id = 123language = "en"status = "published"
Filtering first is usually better, because it reduces the search space and improves relevance within that slice.
If your vector DB supports indexes on metadata
Make sure metadata fields used in filters are indexed if possible. That improves performance.
If your DB uses approximate nearest neighbors (ANN)
Filtering may be implemented as:
- pre-filtering: search only within filtered candidates
- post-filtering: search broadly, then apply filters
- integrated filtering: the engine handles both together
Integrated filtering is usually most efficient.
Example with a typical vector database API
results = collection.query(
vector=query_embedding,
top_k=5,
filter={
"doc_type": "report",
"published": True
}
)
This returns the 5 nearest vectors among documents where:
doc_type == "report"published == True
Rule of thumb
- Use metadata filters for hard constraints: tenant, permissions, language, date range, category
- Use vector similarity for semantic ranking
- Prefer a database/query engine that supports both in one query
If you tell me which vector database or search stack you’re using (Pinecone, Weaviate, Chroma, Milvus, Elasticsearch, pgvector, etc.), I can show the exact syntax.