Prompt
how to index embeddings in qdrant with filters
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
In Qdrant, you usually don’t “index embeddings with filters” as one combined thing. Instead, you:
- Store embeddings in a vector field
- Store metadata in payload fields
- Create payload indexes on the fields you want to filter by
- Search with a vector + filter at query time
1) Create a collection with vectors
Example using sentence-transformers style vectors:
from qdrant_client import QdrantClient
from qdrant_client.models import VectorParams, Distance
client = QdrantClient(url="http://localhost:6333")
client.create_collection(
collection_name="docs",
vectors_config=VectorParams(size=384, distance=Distance.COSINE),
)
2) Insert vectors with payload metadata
Put your filterable fields in payload:
client.upsert(
collection_name="docs",
points=[
{
"id": 1,
"vector": [0.1] * 384,
"payload": {"category": "news", "lang": "en", "year": 2024},
},
{
"id": 2,
"vector": [0.2] * 384,
"payload": {"category": "blog", "lang": "en", "year": 2023},
},
],
)
3) Create payload indexes for filterable fields
This helps Qdrant filter faster:
from qdrant_client.models import PayloadSchemaType
client.create_payload_index(
collection_name="docs",
field_name="category",
field_schema=PayloadSchemaType.KEYWORD,
)
client.create_payload_index(
collection_name="docs",
field_name="lang",
field_schema=PayloadSchemaType.KEYWORD,
)
client.create_payload_index(
collection_name="docs",
field_name="year",
field_schema=PayloadSchemaType.INTEGER,
)
Common payload index types
KEYWORDfor exact-match stringsINTEGER/FLOATfor numeric rangesBOOLDATETIMEGEO
4) Search with a vector and a filter
Example: only search English news from 2024:
from qdrant_client.models import Filter, FieldCondition, MatchValue, Range, SearchParams
query_vector = [0.1] * 384
results = client.search(
collection_name="docs",
query_vector=query_vector,
limit=5,
query_filter=Filter(
must=[
FieldCondition(key="category", match=MatchValue(value="news")),
FieldCondition(key="lang", match=MatchValue(value="en")),
FieldCondition(key="year", range=Range(gte=2024)),
]
),
)
Important note
Qdrant uses the vector index for similarity search and the payload index for filtering.
So “index embeddings with filters” really means:
- Vector index: for nearest-neighbor search
- Payload index: for metadata filters
Optional: filter before or after search?
Qdrant can combine them efficiently. Depending on the dataset and filter selectivity, it may:
- filter first, then search a smaller subset
- search candidates, then apply filter
- use optimized hybrid logic internally
Best practices
- Index fields you filter on often
- Use exact types (
KEYWORD,INTEGER, etc.) - Avoid storing filterable data only inside the text blob
- Keep payload small and structured
If you want, I can also show:
- a REST API example
- a LangChain example
- how to do hybrid search + filters in Qdrant