Prompt
What should I use for embeddings if I need both text and image search?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If you need both text and image search, the best default is to use a multimodal embedding model that maps text and images into the same vector space.
Good choices
1) OpenAI text-embedding-3-large + image embeddings from a multimodal model
- Best if you want to stay in the OpenAI ecosystem and your use case is mostly text.
- But note: OpenAI’s text embedding models are text-only, so for image search you’d need a separate image embedding model and then handle cross-modal alignment yourself.
2) CLIP / OpenCLIP
- Classic and widely used for text-image joint search.
- You can embed:
- text queries
- images
- compare them directly in the same vector space
- Great for:
- “find images similar to this caption”
- “search images with a text query”
- Strong practical default if you need one model for both modalities.
3) SigLIP / SigLIP2
- Often a strong alternative to CLIP.
- Better retrieval quality in many cases.
- Good choice if you want modern multimodal retrieval performance.
4) Commercial multimodal APIs
If you want the easiest production path, look for APIs that explicitly support cross-modal embeddings or multimodal retrieval. These can save engineering time, but check:
- whether text and images land in the same embedding space
- latency/cost
- batch support
- resolution limits for images
Recommendation by scenario
- One embedding space for both text and image search: use CLIP/OpenCLIP or SigLIP
- Best text search only: use a strong text embedding model like
text-embedding-3-large - Best image understanding + search in production: use a multimodal model designed for retrieval, not separate text/image models unless you plan to align them yourself
Practical tip
For many systems, a strong setup is:
- Use CLIP/SigLIP embeddings for cross-modal retrieval
- Optionally add a reranker:
- text reranker for text queries
- vision-language reranker for image search
- Store metadata separately for filtering
Bottom line
If you need both text and image search in the same index, start with CLIP or SigLIP.
If you want, I can also suggest:
- open-source models by quality/speed
- hosted APIs
- how to build the index schema for mixed text/image search