Prompt
what is the best way to search embeddings with metadata
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
The best way is usually:
- Store embeddings in a vector database
- Examples: Pinecone, Weaviate, Milvus, Qdrant, Elasticsearch/OpenSearch, Postgres + pgvector.
- Store metadata alongside each vector
- Things like
type,date,author,tenant_id,category,language, etc.
- Things like
- Use metadata filtering before or during vector search
- Example: “search only documents where
tenant_id = 123andcategory = legal.”
- Example: “search only documents where
- Then rank by vector similarity
- Typically cosine similarity, dot product, or L2 depending on your embeddings.
Best practice
Use a system that supports hybrid retrieval:
- Metadata filters for hard constraints
- Vector similarity for semantic matching
- Optionally keyword/BM25 search for exact terms
This is better than searching embeddings alone because metadata helps narrow the search space and improves precision.
Common patterns
1. Filter first, then vector search
Good when metadata is strict:
- tenant, permissions, language, document type, date range
Example:
- “Find similar passages, but only in English and only for customer A.”
2. Vector search first, then filter/rerank
Useful when metadata is loose or the vector index is small.
3. Hybrid search
Best for most production systems:
- combine semantic score + keyword score + metadata constraints
Example query
“Find contracts similar to this clause, only from 2024, only for legal docs.”
You’d do:
- metadata filter:
year = 2024 AND doc_type = contract AND department = legal - vector search: similarity against the clause embedding
If you’re building from scratch
A good setup is:
- Postgres + pgvector for simpler systems
- Qdrant/Weaviate/Pinecone for larger-scale or more advanced filtering
- Elasticsearch/OpenSearch if you need strong text search + vector + filters
Rule of thumb
- Small/simple: Postgres + pgvector
- Medium/large with rich metadata filters: Qdrant or Weaviate
- Search-heavy enterprise: Elasticsearch/OpenSearch
- Managed scaling: Pinecone
If you want, I can also recommend the best approach for your specific use case (RAG, product search, document search, multi-tenant app, etc.).