Prompt
What do I use for search over text chunks and images together?
Latest observation
Use a multimodal search / vector database setup.
For text chunks + images together, the most common approach is:
-
Embed both modalities into the same vector space
- Text: use a text embedding model
- Images: use a multimodal/image-text embedding model like CLIP or a similar model
-
Store the embeddings in a vector database
- Examples: Pinecone, Weaviate, Milvus, Qdrant, Elasticsearch/OpenSearch with vector search
-
Search with a text query or image query
- Convert the query to an embedding
- Retrieve the nearest neighbors across both text and images
Good options depending on your goal
-
Best general answer:
CLIP + vector database -
If you want one system for text, images, and metadata:
Weaviate or Qdrant with multimodal embeddings -
If you need enterprise search with hybrid keyword + vector:
Elasticsearch/OpenSearch -
If you want simple developer-managed vector search:
Pinecone or Qdrant
Important note
If you want cross-modal search like:
- “find images related to this paragraph”
- “find text describing this image”
then you need a shared embedding model such as CLIP.
If you only embed text and images separately with unrelated models, they won’t be directly comparable.
If you want, I can recommend a specific stack based on whether you’re building in Python, JavaScript, or a cloud platform.