Prompt
I’m building image similarity search and need an embeddings stack that scales
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For image similarity search at scale, a solid embeddings stack usually has these layers:
1) Image embedding model
Pick a model that outputs a fixed-length vector for each image.
Good options:
- CLIP / OpenCLIP: strong general-purpose image embeddings, especially if you may also want text-to-image search later.
- DINOv2: very strong visual features for pure image similarity.
- SigLIP: good alternative to CLIP with strong retrieval performance.
Typical embedding sizes:
- 512, 768, or 1024 dimensions
Recommendation:
- If you want image + text search: OpenCLIP
- If you want image-only similarity: DINOv2
2) Preprocessing pipeline
To keep embeddings consistent:
- resize/crop to model’s expected input size
- normalize with model-specific mean/std
- deduplicate near-identical inputs if needed
- store image IDs and metadata alongside vectors
3) Vector storage + index
At scale, you’ll want an ANN index rather than brute-force search.
Good choices:
- FAISS: best for custom/self-hosted high-performance search
- Milvus: good distributed vector DB
- Qdrant: easy to use, strong filtering support
- Weaviate: useful if you want hybrid search and metadata filtering
Index types:
- HNSW: great recall/latency tradeoff, easy to use
- IVF-PQ: better for very large datasets with memory constraints
- Flat: only for small datasets or exact search
Recommended default:
- HNSW for up to millions of images
- IVF-PQ if you’re going to tens/hundreds of millions and need compression
4) Serving architecture
A scalable setup usually looks like:
- batch embedding service for ingestion
- online embedding service for queries
- vector DB / ANN index for retrieval
- metadata store for image info, permissions, and filters
Common pattern:
- ingest image
- generate embedding
- write vector + metadata
- query image comes in
- embed query
- ANN search top-k
- optional reranking
5) Reranking
For better precision, use a second-stage reranker:
- compute cosine similarity on top-k candidates
- or use a stronger cross-modal / pairwise similarity model
- optionally add business rules and metadata filters
6) Scaling considerations
If you expect large scale:
- use batch inference on GPUs for throughput
- quantize embeddings or use PQ for memory savings
- shard by index or tenant if needed
- cache frequent query embeddings/results
- monitor recall, latency, and drift
A practical stack I’d recommend
If you want fastest path to production:
- Model: OpenCLIP or DINOv2
- Embedding service: PyTorch + GPU batch worker
- Index: Qdrant or Milvus for managed vector search
- Storage: Postgres/S3 for metadata and raw images
- Rerank: cosine similarity on top 100 candidates
If you want maximum control:
- Model: OpenCLIP/DINOv2 in PyTorch
- Index: FAISS HNSW or IVF-PQ
- Metadata: Postgres
- Orchestration: Kafka/SQS + worker pool
- Serving: FastAPI/gRPC
Rule of thumb
- < 1M images: HNSW is usually simplest
- 1M–50M: HNSW or IVF depending on memory
- 50M+: IVF-PQ, sharding, compression, and batch-heavy ingestion
If you want, I can give you:
- a reference architecture
- a model comparison matrix
- a production-ready stack recommendation based on your scale/latency budget