Prompt
How do I search across text embeddings with metadata filters?
Latest observation
To search across text embeddings with metadata filters, you typically do a vector similarity search plus a structured metadata filter.
Basic idea
-
Store each text chunk/document with:
- its embedding vector
- metadata like
author,date,category,source,tenant_id, etc.
-
When querying:
- convert the user query into an embedding
- search for the nearest vectors
- apply metadata conditions to restrict the candidate set
Common patterns
1) Filter first, then vector search
Use metadata to narrow the dataset, then compute similarity only on matching records.
Example:
- “Find similar docs from 2024 in finance”
- Filter:
year = 2024 AND category = 'finance' - Then do vector search within that subset
This is common when the metadata is selective and can reduce cost.
2) Vector search first, then filter
Retrieve top-k similar vectors, then discard those that don’t match metadata.
This is simpler but can be less accurate if the filter is restrictive, because good matches may be excluded after retrieval.
3) Combined hybrid query
Many vector databases support both in one query:
- similarity ranking by embedding distance
- metadata predicate in the same request
This is usually the best option.
Example query logic
Suppose each record looks like:
{
"id": "doc_123",
"text": "How to file expense reports",
"embedding": [0.12, -0.44, ...],
"metadata": {
"department": "finance",
"created_at": "2024-03-12",
"language": "en"
}
}
Query:
- “expense policy”
- filter:
department = finance AND language = en
The system:
- embeds
"expense policy" - finds nearest embeddings
- keeps only items with matching metadata
- returns top results
Example in pseudocode
query_vec = embed("expense policy")
results = vector_search(
vector=query_vec,
top_k=10,
filter={
"department": "finance",
"language": "en"
}
)
Example SQL-like filter
If your system supports SQL-style predicates:
SELECT id, text
FROM documents
WHERE department = 'finance'
AND created_at >= '2024-01-01'
ORDER BY embedding <-> :query_embedding
LIMIT 10;
Tips for good results
- Index metadata fields you filter on often.
- Use high-cardinality filters carefully; they can strongly reduce candidates.
- Keep embeddings and metadata in the same record or strongly linked.
- If using chunks, store metadata from the parent document on each chunk.
- Consider hybrid search (vector + keyword) if exact terms matter.
If you're using a vector DB
Different systems use slightly different syntax:
- Pinecone: metadata filter in the query
- Weaviate:
wherefilters - Milvus: boolean expressions
- Qdrant: payload filters
- pgvector/Postgres: combine vector distance with
WHERE - Elasticsearch/OpenSearch: kNN with filters or hybrid retrieval
If you want, I can show:
- a concrete example in Pinecone, Qdrant, Weaviate, or pgvector, or
- a recommended schema for storing embeddings + metadata.