Prompt
need semantic search for documents with date filters
Latest observation
Yes — you can build semantic search over documents with date filters. The usual pattern is:
- Ingest documents
- Extract text + metadata (e.g.
created_at,published_at,source,author) - Create embeddings for the text
- Store embeddings + metadata in a vector database or search engine
- At query time:
- run semantic similarity search
- apply date range filters on metadata
Common architectures
1) Vector DB with metadata filtering
Good if you want semantic search first, then filter by dates.
Examples:
- Pinecone
- Weaviate
- Qdrant
- Milvus
- Chroma (smaller setups)
Typical query:
- similarity: “find documents about contract renewal”
- filter:
date >= 2024-01-01 AND date <= 2024-12-31
2) Search engine + vectors
Good if you also need strong keyword search and faceting.
Examples:
- Elasticsearch / OpenSearch
- Azure AI Search
These support:
- vector similarity
- text search
- structured filters like dates
3) Hybrid search
Best for quality in many cases:
- semantic/vector ranking
- keyword match
- date filters
- reranking
Recommended data model
Store each document chunk with metadata like:
{
"id": "doc_123_chunk_4",
"text": "....",
"embedding": [0.12, -0.07, ...],
"metadata": {
"document_id": "doc_123",
"title": "Quarterly report",
"created_at": "2025-01-15",
"updated_at": "2025-01-20",
"source": "sharepoint"
}
}
If documents are long, chunk them and keep the document date on every chunk.
Example query logic
User asks:
“Find documents about vendor risk from last quarter”
Translate to:
- semantic query:
"vendor risk" - date filter: last quarter range
Example:
2025-04-01to2025-06-30
Then search only documents whose metadata date falls in that range.
Important design choices
Which date to filter on?
Decide whether to use:
created_atpublished_atmodified_atdocument_date(business date)
This matters a lot. Usually you should expose one canonical date field for filtering.
Filter before or after semantic search?
Usually:
- apply date filter at retrieval time if your database supports it
- if not, retrieve top-N semantically then post-filter, but that can hurt recall
Chunk-level or document-level filtering?
If documents are chunked:
- keep metadata on each chunk
- filter chunks by document date
If you want a simple stack
Small/medium scale
- Python
- embeddings from OpenAI / sentence-transformers
- Qdrant or Chroma
- date metadata filters
Enterprise/search-heavy
- OpenSearch or Elasticsearch
- vectors + BM25 + date facets
Example with Qdrant-style filter
Conceptually:
search(
query_vector=embedding("vendor risk"),
filter={
"must": [
{"field": "created_at", "gte": "2025-04-01"},
{"field": "created_at", "lte": "2025-06-30"}
]
}
)
If you want, I can help with any of these next:
- a system design
- a Python implementation
- choosing between Pinecone / Qdrant / Elasticsearch
- adding natural-language date parsing like “last month”, “Q2 2025”
If you tell me your stack and document volume, I can recommend the best approach.