Prompt
How do I create embeddings for images and search them by similarity?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To create embeddings for images and search them by similarity, the usual workflow is:
- Convert each image into a vector using an image embedding model.
- Store those vectors in a vector database or index.
- For a new query image, embed it the same way.
- Compare vectors with a similarity metric like cosine similarity or dot product.
- Return the nearest vectors/images.
1) Choose an embedding model
Common options:
- CLIP: great if you want image-to-image and text-to-image search.
- ResNet / EfficientNet / ViT: good for image similarity.
- OpenAI image embedding models if you want a managed API.
- Sentence Transformers CLIP variants for easy local use.
If you want search by both image and text, CLIP is usually the best starting point.
2) Generate embeddings
Example with CLIP in Python
import torch
from PIL import Image
import open_clip
# Load model
model, preprocess, tokenizer = open_clip.create_model_and_transforms(
'ViT-B-32', pretrained='openai'
)
model.eval()
# Load image
image = preprocess(Image.open("image.jpg")).unsqueeze(0)
# Get embedding
with torch.no_grad():
image_embedding = model.encode_image(image)
# Normalize for cosine similarity
image_embedding = image_embedding / image_embedding.norm(dim=-1, keepdim=True)
print(image_embedding.shape)
This gives you a vector representation of the image.
3) Build an index of embeddings
Store each image’s vector along with metadata such as:
- image path or URL
- tags
- timestamp
- category
Simple NumPy example
import numpy as np
embeddings = np.array([
# vector1, vector2, ...
])
# Save to disk
np.save("image_embeddings.npy", embeddings)
For anything larger than a few thousand images, use a vector index like:
- FAISS
- Milvus
- Pinecone
- Weaviate
- Qdrant
- Chroma
4) Search by similarity
Cosine similarity manually
from sklearn.metrics.pairwise import cosine_similarity
query_embedding = image_embedding.cpu().numpy()
scores = cosine_similarity(query_embedding, embeddings)[0]
top_k = scores.argsort()[::-1][:5]
print(top_k)
5) Example with FAISS
FAISS is a popular local vector search library.
import faiss
import numpy as np
# embeddings: shape (N, D), float32
embeddings = embeddings.astype("float32")
# Cosine similarity usually uses normalized vectors + inner product
index = faiss.IndexFlatIP(embeddings.shape[1])
index.add(embeddings)
# Query
query = query_embedding.astype("float32")
scores, indices = index.search(query, 5)
print(indices[0], scores[0])
If you normalize embeddings first, inner product becomes cosine similarity.
6) Typical production setup
A practical pipeline looks like this:
-
Ingestion
- Load image
- Preprocess
- Generate embedding
- Save to vector DB
-
Query
- Embed query image
- Run nearest-neighbor search
- Return matching image IDs/paths
7) Tips for good results
- Use the same preprocessing for all images.
- Normalize embeddings if using cosine similarity.
- Choose a model trained on general images if your dataset is diverse.
- If you need domain-specific search, fine-tune or use a model closer to your domain.
- Store metadata so you can filter results after retrieval.
8) If you want text search too
With CLIP, you can also search images using text:
text = tokenizer(["a red sports car"])
with torch.no_grad():
text_embedding = model.encode_text(text)
Then compare text embedding to image embeddings in the same vector space.
If you want, I can give you:
- a complete working Python script for image similarity search,
- a FAISS + CLIP example, or
- a vector database example with Qdrant/Pinecone/Weaviate.