Prompt
I’m building a multimodal search app with text and images; what embedding workflow makes sense?
Latest observation
A good multimodal embedding workflow depends on what you want the search experience to feel like:
- Text → text
- Image → image
- Text → image
- Image → text
- or all of the above in one unified search space
For most apps, the best practical choice is a shared embedding space using a model like CLIP/SigLIP-style embeddings, because then text and images can be compared directly with cosine similarity / dot product.
Recommended workflow
1) Pick your retrieval mode
Most multimodal apps use one of these:
A. Unified cross-modal search
Use if users search with either text or images and want results of both types.
- Embed text and images into the same vector space
- Store one vector per item
- Search queries can be text or image
Good for:
- product search
- media libraries
- e-commerce
- visual discovery
B. Separate embeddings + fusion
Use if you need stronger performance on each modality and can tolerate more complexity.
- One text embedding model
- One image embedding model
- Optionally fuse scores from multiple indexes
Good for:
- large catalogs
- domain-specific retrieval
- systems needing more control
C. Multistage retrieval
Use if relevance matters most.
- First-stage vector retrieval
- Then rerank with a cross-encoder / multimodal model
Good for:
- high-precision search
- enterprise search
- content moderation / compliance workflows
Best-practice architecture
Ingestion pipeline
For each document/item:
-
Extract text
- title
- caption
- OCR text
- metadata
- user tags
-
Process images
- resize/normalize
- optionally generate thumbnails
- optionally extract OCR or region crops
-
Create embeddings
text_embedding = embed(text)image_embedding = embed(image)- if using a shared space, both use the same model family
-
Store
- vector(s)
- metadata
- original content pointers
- modality type
-
Index
- vector database / ANN index
- metadata filters
- optional lexical index (BM25) alongside vectors
Query workflow
Text query
- Embed the query text
- Retrieve nearest neighbors from the vector index
- Optionally rerank with:
- lexical match
- business rules
- multimodal reranker
- Return ranked results
Image query
- Embed the query image
- Retrieve nearest neighbors from the same vector index
- Optionally rerank
- Return results
Hybrid query
If query includes both text and image:
- embed both
- either average/fuse embeddings
- or retrieve separately and merge scores
Model choice
If you want one model for both text and images
Use a CLIP-like model.
Benefits:
- simple
- works well for cross-modal retrieval
- one embedding space
- easy operationally
Tradeoff:
- may be less optimal than specialized models for some domains
If your domain is specialized
Consider:
- domain-tuned CLIP variants
- fine-tuning on your own image-text pairs
- OCR-aware models if documents/screenshots matter
Examples of useful data:
- image + caption
- product image + title + description
- screenshot + extracted text
- video frame + transcript segment
What to store in the index
For each item, store:
idmodality= image/text/bothembeddingor embeddingstext fieldsfor BM25 / filtersimage URLor blob pointersourcetagstimestamplanguagepermissions / tenant info
If using multiple embeddings per item, common patterns are:
- one embedding for title
- one for body text
- one for image
- one fused “item” embedding
Important design decisions
1) Single vector vs multiple vectors per item
Single vector
- easier
- cheaper
- less control
Multiple vectors
- better recall
- handles long documents and multiple images better
- more storage and complexity
A good default:
- one vector for each important field
- retrieve top-k per field
- merge results
2) Text chunking
For long text, don’t embed the whole thing as one vector.
Instead:
- split into chunks of ~200–500 tokens
- embed each chunk
- store chunk-level vectors
- optionally also store a document-level summary embedding
3) OCR and captions matter
If you search screenshots, receipts, slides, or PDFs:
- run OCR
- embed OCR text too
- optionally combine OCR + caption + surrounding metadata
This often boosts retrieval more than image embeddings alone.
4) Add lexical search
Vector search is strong semantically, but lexical search is still valuable.
Best setup:
- vector search for semantic recall
- BM25 / keyword index for exact matches
- hybrid ranking to combine them
This is especially helpful for:
- names
- SKUs
- technical terms
- product IDs
- exact phrases
A practical default setup
If you want a simple but strong baseline:
- Model: CLIP/SigLIP-style shared embedding model
- Index: vector DB supporting ANN + metadata filters
- Text augmentation: OCR, captions, titles, tags
- Hybrid search: vector + BM25
- Reranking: optional cross-encoder for top 50–200 candidates
- Chunking: chunk long text into passages
- Fusing: if item has both text and image, store both and optionally a fused embedding
Example retrieval strategy
For an item with text and image
Store:
title_embeddingbody_chunk_embeddingsimage_embedding
At query time:
- Embed query
- Search all fields
- Combine scores:
- title match weighted highest
- image match moderate
- body match moderate
- Rerank top results
If you’re early-stage
Start with:
- one shared multimodal embedding model
- text + image embeddings in same space
- hybrid BM25 + vector
- metadata filters
- reranking later
That gives you a fast path to a working product without overengineering.
Rule of thumb
- Need easy cross-modal search? Use a shared embedding model.
- Need best relevance? Use shared embeddings + reranker.
- Need domain precision? Fine-tune on your own pairs and add OCR/metadata.
- Need exact matches too? Add BM25.
If you want, I can also sketch:
- a reference architecture,
- a schema for storing embeddings, or
- a model shortlist for your specific use case.