Prompt
I need a practical recommendation for storing embeddings from text, images, and product data in one system.
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
A practical setup is:
Recommended approach
Use one vector database for all embeddings, with separate namespaces/collections or a shared collection with metadata for:
- text embeddings
- image embeddings
- product/data embeddings
This gives you:
- one retrieval system
- unified search across modalities
- simpler ops than running separate stores
Best-practice design
1) Store each item with:
idembeddingmodality=text | image | productsource_idordocument_idmetadatasuch as:- title
- tags
- category
- timestamp
- language
- product attributes
- image URL / file reference
2) Keep raw data outside the vector DB
Store the original:
- text in your document store / object storage
- images in S3/GCS/Azure Blob
- product records in SQL/NoSQL
The vector DB should hold the embedding + lookup metadata, not be your system of record.
3) Use a schema that supports filtering
Example filterable fields:
modalitybrandcategorylanguageprice_rangeavailabilitytenant_id
That lets you do:
- “search only images”
- “search only products in electronics”
- “search text and products together”
Strong practical choices
If you want the simplest production option:
- Pinecone, Weaviate, or Qdrant
- all support metadata filtering and scalable vector search
If you want one open-source stack:
- Qdrant is a very practical default
- easy to run
- good metadata filtering
- works well for multimodal embeddings
If you already use Postgres:
- pgvector can work well for smaller to medium systems
- especially if your scale is modest
- but it’s usually less convenient for heavy multimodal search at large scale
Important implementation detail
If your text, image, and product embeddings come from different models, that’s fine, but:
- keep them in separate collections if dimensions or distance metrics differ
- or store them in one collection only if they are comparable and you normalize the schema carefully
In practice, separate collections are often cleaner:
text_vectorsimage_vectorsproduct_vectors
Then search across them at the application layer if needed.
My recommendation
If you want a balanced, practical choice:
Use Qdrant with separate collections by modality and shared metadata fields.
- raw text/images/products stay in their own storage
- embeddings go into Qdrant
- your app queries one or more collections depending on the use case
If you want, I can give you:
- a concrete schema example, or
- a decision matrix comparing Qdrant vs Pinecone vs pgvector for your case.