Prompt
How do I build semantic search with embeddings and metadata filters?
Latest observation
To build semantic search with embeddings + metadata filters, think of it as a two-stage system:
- Find semantically similar items using vector embeddings
- Restrict or rerank results using structured metadata filters
1) Core idea
Each document/item gets:
- a text embedding: a vector representation of its meaning
- metadata: structured fields like:
categoryauthorcreated_atlanguagetenant_idpricetags
At query time:
- embed the user query
- compare the query vector to stored vectors
- apply metadata filters such as:
category = "support"created_at >= 2024-01-01tenant_id = 123
2) Typical architecture
Ingestion
- Collect documents
- Split long docs into chunks
- Generate embeddings for each chunk
- Store:
- chunk text
- embedding vector
- metadata
- document ID / chunk ID
Query
- User enters a search query
- Embed the query
- Search vector index for nearest neighbors
- Apply metadata filters
- Return top results
3) Data model
A record might look like:
{
"id": "chunk_123",
"text": "How to reset your password...",
"embedding": [0.12, -0.44, 0.98, ...],
"metadata": {
"doc_id": "doc_45",
"category": "help_center",
"language": "en",
"tenant_id": "acme",
"created_at": "2025-01-10"
}
}
4) How filtering works
There are two common approaches:
A. Pre-filter then vector search
Apply metadata constraints first, then run vector similarity only on matching items.
Pros
- precise filtering
- good for strong constraints like
tenant_id,language
Cons
- may reduce recall if too restrictive
- can be slower if your filter set is large and not well indexed
B. Vector search then post-filter
Find nearest neighbors first, then filter the top candidates.
Pros
- simpler to implement
- can work when filters are weak
Cons
- might miss valid matches that were filtered out after retrieval
Best practice
Use a vector database / search engine that supports hybrid filtering, so filtering happens efficiently during retrieval.
5) Technologies you can use
Common options:
- Postgres + pgvector
- Pinecone
- Weaviate
- Qdrant
- Milvus
- Elasticsearch / OpenSearch with vector support
- FAISS plus your own metadata store
If you need metadata filters, choose a system that supports them natively.
6) Example with pgvector
Table schema
CREATE TABLE documents (
id bigserial PRIMARY KEY,
content text,
embedding vector(1536),
category text,
language text,
tenant_id text,
created_at timestamp
);
Query with filter + similarity
SELECT id, content
FROM documents
WHERE category = 'help_center'
AND language = 'en'
ORDER BY embedding <-> '[0.1, 0.2, ...]'::vector
LIMIT 10;
Here:
embedding <-> query_vector= distance operatorWHEREclause applies metadata filters
For performance, add indexes on metadata columns and a vector index on embeddings.
7) Example with a vector DB filter
In Pinecone-style pseudocode:
results = index.query(
vector=query_embedding,
top_k=10,
filter={
"category": {"$eq": "help_center"},
"language": {"$eq": "en"},
"tenant_id": {"$eq": "acme"}
}
)
8) Hybrid search: semantic + keyword
Semantic search is often better when combined with keyword matching.
A strong pattern is:
- lexical score from BM25 / full-text search
- semantic score from embeddings
- metadata filter for narrowing results
This helps when:
- exact product names matter
- rare terms matter
- users search with short queries
9) Chunking matters
For long documents, split into chunks like:
- 200–500 tokens per chunk
- slight overlap, e.g. 10–20%
Store chunk metadata:
- document ID
- section title
- page number
- chunk order
This improves retrieval precision.
10) Practical design tips
Use metadata for hard constraints
Examples:
- tenant isolation
- language
- access permissions
- date range
- content type
Use embeddings for soft meaning
Examples:
- “how do I change my password?”
- “reset login credentials”
- “account access issue”
Keep metadata clean and typed
Use consistent field values:
language: "en"- not sometimes
"english"and sometimes"EN"
Index filter fields
If a field will be filtered often, index it.
11) Common pitfalls
- Filtering after retrieval only can miss relevant results
- Poor chunking hurts semantic quality
- No metadata normalization makes filters unreliable
- Using embeddings alone for exact-match queries can be weak
- Not separating tenant data can cause security issues
12) Recommended implementation pattern
If you’re starting from scratch:
Simple stack
- Postgres
- pgvector
- full-text search
- application-side metadata handling
Larger scale
- Qdrant / Pinecone / Weaviate
- embedding pipeline
- metadata filters in the vector index
- reranking model for top results
13) Minimal workflow
- Embed documents
- Store embeddings + metadata
- On query, embed query
- Search by vector similarity
- Apply filters
- Return top-k
- Optional: rerank results
14) If you want the best quality
Use a 3-step retrieval pipeline:
- Filter by permissions / tenant / language
- Retrieve top 50–100 with vector similarity
- Rerank top candidates with a cross-encoder or LLM
This usually gives better quality than vector search alone.
If you want, I can also show you:
- a Python example with pgvector
- a Pinecone/Qdrant example
- or a reference architecture for production semantic search