Prompt
How do I search images and text chunks in the same embeddings database?
Latest observation
To search images and text chunks in the same embeddings database, the key idea is to store them in a shared vector space and give each record metadata that tells you what modality it came from.
1) Use one embedding model or one aligned embedding space
You need embeddings that are comparable across both text and images:
- Best option: use a multimodal model that embeds both images and text into the same vector space.
- Examples: CLIP, SigLIP, OpenCLIP
- If you use separate models for text and images, their vectors usually won’t be directly searchable together unless you project them into a shared space.
2) Store both in the same vector database
Each record can look like this:
{
"id": "img_001",
"vector": [0.12, -0.44, ...],
"modality": "image",
"source": "product_photo.jpg",
"caption": "red running shoe on white background"
}
{
"id": "txt_123",
"vector": [0.08, -0.40, ...],
"modality": "text",
"source": "manual.pdf#chunk_7",
"text": "The shoe is designed for trail running..."
}
3) Query with whichever modality you want
Text query
Embed the text query using the text encoder from your multimodal model:
query_vec = embed_text("red running shoes")
results = vector_db.search(query_vec, top_k=10)
This can return both:
- relevant images
- relevant text chunks
Image query
Embed the image using the image encoder:
query_vec = embed_image("shoe_photo.jpg")
results = vector_db.search(query_vec, top_k=10)
Again, results can include both modalities.
4) Add metadata filters if needed
You may want to search across both, or only one type.
Examples:
- search only images:
filter modality = "image" - search only text:
filter modality = "text" - search both: no filter
5) Good design patterns
Pattern A: Unified embeddings + one index
Simplest and most common.
- One vector index
- Mixed records
- Metadata for modality
- One query returns everything
Pattern B: Separate indexes + merge results
Useful if you have different vector sizes/models, but then you need:
- one text index
- one image index
- query both
- normalize scores
- merge/rerank results
This is more complex and usually only needed if you can’t use a shared embedding space.
6) Important caveats
- Vector similarity scores may not be perfectly calibrated across modalities, even in shared spaces.
- Image embeddings may work better with captions or OCR text stored as metadata.
- For best search quality, combine:
- vector search
- metadata filters
- optional reranking with a cross-encoder or LLM
7) Practical architecture
A robust setup often looks like:
-
Ingest text chunks
- chunk documents
- embed each chunk
- store in vector DB with metadata
-
Ingest images
- embed image
- optionally generate caption/OCR
- store in same vector DB with metadata
-
Search
- embed query text or image
- search single index
- filter/rerank by modality if needed
8) Example with pseudo-schema
id
vector
modality -- text | image
content -- text chunk or image caption
uri -- file path or object storage URL
doc_id
page_num
tags
9) Example query flow
User asks: “show me diagrams about vector databases”
- Embed query text
- Search all vectors
- Return:
- diagrams/images
- text chunks from docs
- Optionally boost records with tags like
diagram,architecture
If you want, I can also show:
- a Python example using CLIP + FAISS/Pinecone/Qdrant, or
- the best schema for storing images + text together in PostgreSQL/pgvector.