Prompt
How do I search images by embedding similarity?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
To search images by embedding similarity, you typically do this:
- Convert each image into a vector embedding using a pretrained model.
- Store those vectors in a vector index / database.
- For a query image, compute its embedding.
- Find nearest neighbors in the index using a similarity metric like cosine similarity or dot product.
Basic workflow
1) Choose an image embedding model
Common options:
- CLIP: great for image-to-image and image-to-text retrieval
- DINOv2: strong visual similarity embeddings
- ResNet / ViT features: usable, but usually less semantically strong than CLIP
2) Extract embeddings
Example with CLIP-like models:
- Input image → model → 512-d or 768-d vector
- Often you L2-normalize embeddings if using cosine similarity
3) Index embeddings
Use a nearest-neighbor library or vector DB:
- FAISS for local/search-at-scale
- Milvus, Pinecone, Weaviate, Qdrant for managed/vector DB solutions
- Annoy / ScaNN for approximate search
4) Query by example
- Embed the query image
- Search top-k nearest vectors
- Return matching images and similarity scores
Similarity metrics
- Cosine similarity: most common for embeddings
- Dot product: often equivalent to cosine if vectors are normalized
- Euclidean distance: also used, but less common for normalized embeddings
Minimal Python example with FAISS
import faiss
import numpy as np
# Suppose embeddings is an (N, D) numpy array of float32
# and query_embedding is shape (D,)
embeddings = np.random.rand(1000, 512).astype("float32")
query_embedding = np.random.rand(512).astype("float32")
# Normalize for cosine similarity
faiss.normalize_L2(embeddings)
faiss.normalize_L2(query_embedding.reshape(1, -1))
# Build index for inner product search
index = faiss.IndexFlatIP(512)
index.add(embeddings)
# Search top 5
scores, ids = index.search(query_embedding.reshape(1, -1), 5)
print(ids[0])
print(scores[0])
Practical considerations
- Precompute embeddings for all images offline
- Store metadata alongside vectors: filename, URL, tags, labels
- Batch embedding extraction for speed
- For large datasets, use approximate nearest neighbor indexing instead of exact search
- If you want semantic similarity, CLIP is usually a strong starting point
- If you want visual similarity (same object, pose, style), DINOv2 or a fine-tuned model may work better
If you want text-to-image search too
Use a multimodal model like CLIP:
- Text query → text embedding
- Image database → image embeddings
- Compare both in the same embedding space
If you want, I can give you:
- a full end-to-end example in Python,
- an example using FAISS + CLIP, or
- a vector database setup like Qdrant or Pinecone.