Prompt
need vector database for images and text together
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
If you want one vector database for both images and text, you’re looking for a multimodal vector DB or a normal vector DB that can store embeddings from different models in the same index/collection.
Good options
1) Pinecone
- Easy managed service
- Works well for text and image embeddings
- Good metadata filtering
- Best if you want low ops overhead
2) Weaviate
- Strong multimodal support
- Can store text, image, and metadata together
- Good for semantic search across modalities
- Open-source + cloud
3) Qdrant
- Very flexible
- Supports named vectors, so you can keep text and image embeddings separately in the same collection
- Great filtering and performance
- Open-source + cloud
4) Milvus
- Scales well
- Supports multimodal use cases
- Good if you need large-scale vector search
- More infra-heavy than Pinecone/Qdrant
5) Postgres + pgvector
- Good for smaller systems
- Can store text/image embeddings in tables
- Simple if you already use PostgreSQL
- Not ideal for very large-scale ANN search compared to dedicated vector DBs
Best approach
For images + text, usually you do this:
- Convert text to text embeddings
- Convert images to image embeddings
- Store them in the same DB with:
- a shared item ID
- modality type (
text,image) - metadata
- Query with the same embedding space if using a multimodal model like CLIP
Important note
If you want search text to find images or image to find text, use a model that maps both into the same vector space, like:
- CLIP
- SigLIP
- multimodal embedding models from OpenAI / others
If you use separate embedding models for text and images, you can still store them together, but cross-modal search is harder unless you align the spaces.
My recommendation
- Fastest to build: Pinecone
- Best open-source multimodal choice: Weaviate or Qdrant
- Best for cross-modal search: use CLIP-style embeddings + one of the above
If you want, I can give you:
- a comparison table
- a recommended stack for your use case
- example code for storing/searching images + text