Prompt

How do I search images and text chunks in the same embeddings database?

Technology / Databases2 observationsLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To search images and text chunks in the same embeddings database, the key idea is to store them in a shared vector space and give each record metadata that tells you what modality it came from.

1) Use one embedding model or one aligned embedding space

You need embeddings that are comparable across both text and images:

  • Best option: use a multimodal model that embeds both images and text into the same vector space.
    • Examples: CLIP, SigLIP, OpenCLIP
  • If you use separate models for text and images, their vectors usually won’t be directly searchable together unless you project them into a shared space.

2) Store both in the same vector database

Each record can look like this:

{
  "id": "img_001",
  "vector": [0.12, -0.44, ...],
  "modality": "image",
  "source": "product_photo.jpg",
  "caption": "red running shoe on white background"
}
{
  "id": "txt_123",
  "vector": [0.08, -0.40, ...],
  "modality": "text",
  "source": "manual.pdf#chunk_7",
  "text": "The shoe is designed for trail running..."
}

3) Query with whichever modality you want

Text query

Embed the text query using the text encoder from your multimodal model:

query_vec = embed_text("red running shoes")
results = vector_db.search(query_vec, top_k=10)

This can return both:

  • relevant images
  • relevant text chunks

Image query

Embed the image using the image encoder:

query_vec = embed_image("shoe_photo.jpg")
results = vector_db.search(query_vec, top_k=10)

Again, results can include both modalities.

4) Add metadata filters if needed

You may want to search across both, or only one type.

Examples:

  • search only images: filter modality = "image"
  • search only text: filter modality = "text"
  • search both: no filter

5) Good design patterns

Pattern A: Unified embeddings + one index

Simplest and most common.

  • One vector index
  • Mixed records
  • Metadata for modality
  • One query returns everything

Pattern B: Separate indexes + merge results

Useful if you have different vector sizes/models, but then you need:

  • one text index
  • one image index
  • query both
  • normalize scores
  • merge/rerank results

This is more complex and usually only needed if you can’t use a shared embedding space.

6) Important caveats

  • Vector similarity scores may not be perfectly calibrated across modalities, even in shared spaces.
  • Image embeddings may work better with captions or OCR text stored as metadata.
  • For best search quality, combine:
    • vector search
    • metadata filters
    • optional reranking with a cross-encoder or LLM

7) Practical architecture

A robust setup often looks like:

  1. Ingest text chunks

    • chunk documents
    • embed each chunk
    • store in vector DB with metadata
  2. Ingest images

    • embed image
    • optionally generate caption/OCR
    • store in same vector DB with metadata
  3. Search

    • embed query text or image
    • search single index
    • filter/rerank by modality if needed

8) Example with pseudo-schema

id
vector
modality   -- text | image
content    -- text chunk or image caption
uri        -- file path or object storage URL
doc_id
page_num
tags

9) Example query flow

User asks: “show me diagrams about vector databases”

  • Embed query text
  • Search all vectors
  • Return:
    • diagrams/images
    • text chunks from docs
  • Optionally boost records with tags like diagram, architecture

If you want, I can also show:

  1. a Python example using CLIP + FAISS/Pinecone/Qdrant, or
  2. the best schema for storing images + text together in PostgreSQL/pgvector.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (2 observations for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.