Prompt

How do I search image embeddings in a database?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

To search image embeddings in a database, you typically do vector similarity search:

1) Convert images to embeddings

Use a pretrained model (for example CLIP, DINOv2, ResNet-based feature extractor, etc.) to turn each image into a fixed-length vector, like:

image -> embedding vector [0.12, -0.44, ...]

Do this for:

  • images you want to index
  • the query image you want to search with

2) Store embeddings in a vector-capable database

Common options:

  • PostgreSQL + pgvector
  • Pinecone
  • Weaviate
  • Milvus
  • Qdrant
  • Elasticsearch / OpenSearch with vector support

Each row/document usually stores:

  • an id
  • the embedding
  • metadata like filename, tags, category, timestamps, etc.

3) Search by similarity

When a query image comes in:

  1. generate its embedding
  2. compare it to stored embeddings using a distance metric:
    • cosine similarity (very common)
    • dot product
    • Euclidean distance
  3. return the nearest neighbors

4) Example: PostgreSQL with pgvector

Schema:

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE images (
  id BIGSERIAL PRIMARY KEY,
  url TEXT,
  embedding vector(512)
);

Insert:

INSERT INTO images (url, embedding)
VALUES ('cat.jpg', '[0.1, 0.2, ...]');

Query nearest images:

SELECT id, url
FROM images
ORDER BY embedding <-> '[0.12, 0.18, ...]'
LIMIT 10;
  • <-> = Euclidean distance in pgvector
  • For cosine distance you can use the appropriate operator/index depending on setup

5) Add an index for scale

Without indexing, searches become slow as data grows.

For pgvector:

CREATE INDEX ON images USING hnsw (embedding vector_cosine_ops);

or:

CREATE INDEX ON images USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);

6) Typical workflow

  • preprocess all images
  • generate embeddings offline
  • insert into DB
  • on query:
    • embed query image
    • run vector similarity search
    • optionally filter by metadata, e.g. category, user_id, date

7) Important tips

  • Use the same embedding model for indexing and querying
  • Normalize embeddings if your similarity metric expects it
  • Store metadata separately for filtering and ranking
  • Rebuild embeddings if you switch models

8) Simple architecture

Image files -> embedding model -> vector DB
Query image -> embedding model -> vector search -> results

If you want, I can also show:

  • a Python example
  • a Postgres/pgvector setup
  • how to do this in Pinecone, Qdrant, or Weaviate

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.