Prompt
what should I use for multimodal embeddings across text and images?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For multimodal embeddings across text and images, the best choice is usually a CLIP-style model or a newer multimodal embedding model trained to put text and images into the same vector space.
Good options
- OpenAI
text-embedding-3-large+ a separate image model
Good for text, but not a single shared space with images. - CLIP / SigLIP / OpenCLIP
Classic choice for shared text-image embeddings. Great for search, retrieval, and matching. - Modern multimodal embedding APIs/models
If you want a managed solution, use a model explicitly designed for cross-modal retrieval.
What to choose depends on your use case
1) Text-to-image / image-to-text retrieval
Use CLIP-like embeddings.
- One encoder for text
- One encoder for images
- Cosine similarity works well
- Best for “find images similar to this text” or “find text matching this image”
2) Semantic search over mixed content
Use a single shared embedding space model that supports both modalities.
- Easier indexing
- Better cross-modal retrieval
- Useful if you want one vector DB for text + images
3) Highest accuracy for specific tasks
You may want:
- a vision-language embedding model for retrieval
- plus a multimodal LLM for reasoning after retrieval
Practical recommendation
If you’re building something today:
- Open-source/local: start with OpenCLIP or SigLIP
- Managed/API: use a multimodal embedding endpoint that explicitly supports both text and images in the same space
- If only text matters most: use a strong text embedding model and add image captions/metadata as text
Rule of thumb
- Need shared text-image search? → CLIP/SigLIP
- Need best text embeddings only? → text embedding model
- Need reasoning over both after retrieval? → embeddings + multimodal LLM
If you want, I can recommend a specific model based on your constraints:
- local vs API
- latency
- budget
- vector DB
- retrieval vs classification vs clustering