Prompt
How do I search image embeddings in a database?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
To search image embeddings in a database, you typically do vector similarity search:
1) Convert images to embeddings
Use a pretrained model (for example CLIP, DINOv2, ResNet-based feature extractor, etc.) to turn each image into a fixed-length vector, like:
image -> embedding vector [0.12, -0.44, ...]
Do this for:
- images you want to index
- the query image you want to search with
2) Store embeddings in a vector-capable database
Common options:
- PostgreSQL + pgvector
- Pinecone
- Weaviate
- Milvus
- Qdrant
- Elasticsearch / OpenSearch with vector support
Each row/document usually stores:
- an
id - the
embedding - metadata like filename, tags, category, timestamps, etc.
3) Search by similarity
When a query image comes in:
- generate its embedding
- compare it to stored embeddings using a distance metric:
- cosine similarity (very common)
- dot product
- Euclidean distance
- return the nearest neighbors
4) Example: PostgreSQL with pgvector
Schema:
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE images (
id BIGSERIAL PRIMARY KEY,
url TEXT,
embedding vector(512)
);
Insert:
INSERT INTO images (url, embedding)
VALUES ('cat.jpg', '[0.1, 0.2, ...]');
Query nearest images:
SELECT id, url
FROM images
ORDER BY embedding <-> '[0.12, 0.18, ...]'
LIMIT 10;
<->= Euclidean distance in pgvector- For cosine distance you can use the appropriate operator/index depending on setup
5) Add an index for scale
Without indexing, searches become slow as data grows.
For pgvector:
CREATE INDEX ON images USING hnsw (embedding vector_cosine_ops);
or:
CREATE INDEX ON images USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);
6) Typical workflow
- preprocess all images
- generate embeddings offline
- insert into DB
- on query:
- embed query image
- run vector similarity search
- optionally filter by metadata, e.g. category, user_id, date
7) Important tips
- Use the same embedding model for indexing and querying
- Normalize embeddings if your similarity metric expects it
- Store metadata separately for filtering and ranking
- Rebuild embeddings if you switch models
8) Simple architecture
Image files -> embedding model -> vector DB
Query image -> embedding model -> vector search -> results
If you want, I can also show:
- a Python example
- a Postgres/pgvector setup
- how to do this in Pinecone, Qdrant, or Weaviate