Prompt
How do I generate image embeddings and search similar images?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To generate image embeddings and search for similar images, the standard approach is:
- Convert each image into a vector embedding
- Store those vectors in a vector index
- For a query image, embed it the same way
- Find nearest vectors by cosine similarity or dot product
1) Generate image embeddings
You can use a pretrained vision model such as:
- CLIP
- DINOv2
- ResNet / EfficientNet feature extractors
- Multimodal embedding models from vector DB providers
Example with CLIP in Python
import torch
from PIL import Image
import clip
device = "cuda" if torch.cuda.is_available() else "cpu"
model, preprocess = clip.load("ViT-B/32", device=device)
image = preprocess(Image.open("example.jpg")).unsqueeze(0).to(device)
with torch.no_grad():
embedding = model.encode_image(image)
# Normalize for cosine similarity
embedding = embedding / embedding.norm(dim=-1, keepdim=True)
print(embedding.shape) # e.g. torch.Size([1, 512])
This gives you a vector representation of the image.
2) Store embeddings in a vector index
For small datasets, you can just keep them in memory and compute similarity directly.
For larger datasets, use a vector database or ANN index such as:
- FAISS
- Pinecone
- Milvus
- Weaviate
- Qdrant
- Chroma
Example with FAISS
import faiss
import numpy as np
# Suppose embeddings is a numpy array of shape (N, D)
embeddings = np.array([...]).astype("float32")
# If using cosine similarity, normalize embeddings first
faiss.normalize_L2(embeddings)
d = embeddings.shape[1]
index = faiss.IndexFlatIP(d) # inner product
index.add(embeddings)
3) Search for similar images
Embed the query image using the same model, then search the index.
query = preprocess(Image.open("query.jpg")).unsqueeze(0).to(device)
with torch.no_grad():
q_emb = model.encode_image(query)
q_emb = q_emb / q_emb.norm(dim=-1, keepdim=True)
q_emb_np = q_emb.cpu().numpy().astype("float32")
# Search top 5
scores, indices = index.search(q_emb_np, 5)
print(indices)
print(scores)
The returned indices point to your stored images, and the scores indicate similarity.
4) Recommended similarity metric
- Cosine similarity is commonly used for image embeddings.
- If embeddings are normalized, cosine similarity becomes equivalent to dot product.
5) Practical tips
- Use the same preprocessing for all images and queries.
- Normalize embeddings if using cosine similarity.
- Store metadata alongside embeddings, such as:
- filename
- URL
- labels
- timestamps
- For large-scale search, use approximate nearest neighbor indices.
6) Simple end-to-end flow
images -> embedding model -> vectors -> vector index
query image -> embedding model -> query vector -> nearest neighbors -> similar images
If you want, I can also give you:
- a complete runnable Python example
- a PyTorch + FAISS pipeline
- or an example using OpenAI / CLIP / Hugging Face models