Prompt
Recommend the best way to generate and store embeddings for text and images at scale
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
The best approach at scale is usually:
1) Use a batched, managed embedding pipeline
- Text embeddings: generate with a strong embedding model (OpenAI, Cohere, or an internal model like bge/e5 depending on your stack).
- Image embeddings: use a vision embedding model (CLIP-like or a multimodal embedding model).
- Run embedding generation in batch jobs or streaming workers, not synchronously in your app path.
Why
- More cost-efficient
- Easier to retry failures
- Lets you scale independently from your user-facing service
2) Store embeddings in two places
A. Canonical object store / data lake
Store:
- source content
- metadata
- model version
- embedding timestamp
- embedding vector
Good options:
- S3 / GCS / Azure Blob
- Parquet / JSONL / Avro for offline processing
This is your source of truth and best for reprocessing.
B. Vector database or ANN index for retrieval
Use a system optimized for similarity search:
- Pinecone
- Weaviate
- Milvus
- Qdrant
- pgvector if your scale is moderate
- FAISS if you manage the index yourself
This is what you query in production.
3) Use a schema that supports versioning
For each embedding record, keep:
idcontent_type(text,image)source_urior raw content referenceembeddingembedding_dimmodel_namemodel_versioncreated_atcontent_hashmetadata(tags, language, product id, etc.)
Important
Never assume embeddings are permanent. If the model changes, you’ll want to regenerate them.
4) Partition by modality and use separate indexes when needed
Text and images often work best with:
- separate embedding models
- separate vector indexes
- or a shared multimodal model only if your use case truly needs cross-modal search
If you need “search by text and retrieve images,” use a shared latent space like CLIP-style embeddings.
5) Batch and chunk strategically
Text
- Split documents into chunks before embedding.
- Common chunk sizes: 200–800 tokens with overlap.
- Store one embedding per chunk, plus document-level metadata.
Images
- Store one embedding per image.
- For documents/PDFs, consider:
- page-level images
- OCR text embeddings
- layout-aware embeddings if needed
6) Design for incremental updates
At scale, avoid re-embedding everything:
- detect content changes with hashes
- embed only new/changed content
- queue reprocessing when model version changes
A simple pattern:
- Ingest content
- Compute content hash
- Check if embedding exists for
(hash, model_version) - If not, enqueue job
- Worker generates embedding
- Persist to object store + vector DB
7) Recommended architecture
Ingestion service
- accepts text/images + metadata
- writes raw content to object storage
- publishes job to queue
Embedding workers
- batch requests to embedding model
- generate vectors
- write results to storage and index
Storage
- S3/GCS for canonical records
- Vector DB for serving retrieval
Query service
- embeds incoming query
- searches vector DB
- optionally reranks with a cross-encoder / multimodal reranker
8) Practical recommendations by scale
Small to medium scale
- Postgres + pgvector
- S3 for raw assets
- background workers
Large scale
- S3 + Kafka/SQS/PubSub + dedicated vector DB
- use batching, sharding, and model versioning
- keep offline parquet snapshots for rebuilds
Very large scale / high throughput
- separate:
- offline embedding pipeline
- online low-latency query pipeline
- consider building ANN indexes with FAISS/HNSW
- maintain replicated indexes and incremental rebuilds
My “best overall” recommendation
If you want the safest scalable default:
- Generate embeddings in async batch workers
- Store canonical embeddings in object storage/data lake
- Index them in a vector database
- Version everything
- Use separate pipelines for text and images unless you need cross-modal retrieval
If you want, I can also recommend a concrete stack for:
- AWS
- GCP
- Azure
- open-source / self-hosted