Prompt
What embedding model should I use for image and text multimodal search?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For multimodal search across image and text, the safest default is:
Recommended default
OpenAI text-embedding-3-large for text + a vision-language model that embeds both image and text into the same space, such as CLIP/SigLIP-style models.
If you want one shared embedding space for both images and text, use a multimodal embedding model rather than a text-only embedding model.
Good options
1) Best general-purpose open-source choice
- SigLIP or CLIP
- Works well for:
- image → text search
- text → image search
- image ↔ image similarity
- Common practical variants:
open_clipimplementations of ViT-L/14- SigLIP variants from Hugging Face
2) Strong commercial/API options
- OpenAI multimodal embeddings if available in your stack
- Google Vertex AI multimodal embeddings
- Cohere multimodal embeddings where supported
These are convenient if you want managed infrastructure and strong out-of-the-box quality.
3) If your use case is mostly text, with images as attached metadata
- Use a text embedding model for text
- Generate image captions / OCR / tags and embed that text too
- This is often surprisingly effective and simpler to operate
- Good when you don’t need true image-to-image semantic matching
Practical recommendation by use case
- Need one model for both images and text: use CLIP/SigLIP
- Need highest quality and can use captions/OCR: use text embeddings + image captioning/OCR
- Need production convenience: use a managed multimodal embedding API
- Need open-source control: use SigLIP or CLIP
What I would choose
If you’re building from scratch and want a solid default:
- SigLIP for the shared image/text embedding space
- Store embeddings in a vector DB
- For images, optionally also extract:
- OCR text
- captions
- metadata
and index those with a text embedding model too
That hybrid setup usually gives the best retrieval quality.
If you want, I can also recommend:
- a specific model based on your deployment constraints,
- the best vector database setup, or
- an architecture for multimodal search.