Prompt
How do I search images using embeddings?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To search images using embeddings, you generally do this:
-
Convert each image into an embedding
- Use an image model to turn every image into a fixed-length vector.
- Similar images end up with vectors that are close together.
-
Store the embeddings in a vector database or index
- Save each vector along with the image ID/path and metadata.
- Common options: FAISS, Pinecone, Milvus, Weaviate, pgvector, Elasticsearch kNN.
-
Embed the query image
- If you want to search by image, run the same embedding model on the query image.
- If you want to search by text, use a multimodal model like CLIP that can embed both text and images into the same vector space.
-
Find nearest neighbors
- Compare the query vector to the stored vectors using cosine similarity, dot product, or Euclidean distance.
- Retrieve the top K most similar images.
-
Optionally re-rank
- Use a more expensive model or additional filters/metadata to refine the results.
Typical workflow
Image-to-image search
- Upload a query image
- Create embedding
- Search for nearest embeddings
- Return visually similar images
Text-to-image search
- Convert text query into embedding with a multimodal model
- Search against image embeddings in the same embedding space
- Return images matching the text concept
Example with CLIP
CLIP is commonly used because it maps images and text into the same vector space.
from PIL import Image
import torch
from transformers import CLIPProcessor, CLIPModel
import numpy as np
import faiss
# Load model
model = CLIPModel.from_pretrained("openai/clip-vit-base-patch32")
processor = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")
# Example images
image_paths = ["img1.jpg", "img2.jpg", "img3.jpg"]
# Build embeddings
embeddings = []
for path in image_paths:
image = Image.open(path).convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
emb = model.get_image_features(**inputs)
emb = emb / emb.norm(dim=-1, keepdim=True) # normalize for cosine similarity
embeddings.append(emb.cpu().numpy())
embeddings = np.vstack(embeddings).astype("float32")
# Build FAISS index
index = faiss.IndexFlatIP(embeddings.shape[1]) # inner product on normalized vectors = cosine similarity
index.add(embeddings)
# Query with text
text = "a red sports car"
inputs = processor(text=[text], return_tensors="pt", padding=True)
with torch.no_grad():
q = model.get_text_features(**inputs)
q = q / q.norm(dim=-1, keepdim=True)
q = q.cpu().numpy().astype("float32")
# Search
scores, idxs = index.search(q, k=3)
print(idxs, scores)
Practical tips
- Normalize embeddings if you plan to use cosine similarity.
- Use a strong pretrained model (CLIP, SigLIP, OpenCLIP, etc.).
- Batch embedding generation for speed.
- Store metadata like filename, labels, timestamps, and categories for filtering.
- Use ANN indexes for large datasets instead of brute-force search.
- Re-index if the model changes.
If you want the simplest architecture
- Use CLIP for embeddings
- Use FAISS for similarity search
- Store metadata in SQLite/Postgres
- Build a small API that:
- accepts an image or text query
- computes embedding
- searches FAISS
- returns top matches
If you want, I can also show:
- a minimal working example
- a production architecture
- or how to do this with pgvector / Pinecone / FAISS