Prompt

I’m building a multimodal search app with text and images; what embedding workflow makes sense?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A good multimodal embedding workflow depends on what you want the search experience to feel like:

  • Text → text
  • Image → image
  • Text → image
  • Image → text
  • or all of the above in one unified search space

For most apps, the best practical choice is a shared embedding space using a model like CLIP/SigLIP-style embeddings, because then text and images can be compared directly with cosine similarity / dot product.

Recommended workflow

1) Pick your retrieval mode

Most multimodal apps use one of these:

A. Unified cross-modal search

Use if users search with either text or images and want results of both types.

  • Embed text and images into the same vector space
  • Store one vector per item
  • Search queries can be text or image

Good for:

  • product search
  • media libraries
  • e-commerce
  • visual discovery

B. Separate embeddings + fusion

Use if you need stronger performance on each modality and can tolerate more complexity.

  • One text embedding model
  • One image embedding model
  • Optionally fuse scores from multiple indexes

Good for:

  • large catalogs
  • domain-specific retrieval
  • systems needing more control

C. Multistage retrieval

Use if relevance matters most.

  • First-stage vector retrieval
  • Then rerank with a cross-encoder / multimodal model

Good for:

  • high-precision search
  • enterprise search
  • content moderation / compliance workflows

Best-practice architecture

Ingestion pipeline

For each document/item:

  1. Extract text

    • title
    • caption
    • OCR text
    • metadata
    • user tags
  2. Process images

    • resize/normalize
    • optionally generate thumbnails
    • optionally extract OCR or region crops
  3. Create embeddings

    • text_embedding = embed(text)
    • image_embedding = embed(image)
    • if using a shared space, both use the same model family
  4. Store

    • vector(s)
    • metadata
    • original content pointers
    • modality type
  5. Index

    • vector database / ANN index
    • metadata filters
    • optional lexical index (BM25) alongside vectors

Query workflow

Text query

  1. Embed the query text
  2. Retrieve nearest neighbors from the vector index
  3. Optionally rerank with:
    • lexical match
    • business rules
    • multimodal reranker
  4. Return ranked results

Image query

  1. Embed the query image
  2. Retrieve nearest neighbors from the same vector index
  3. Optionally rerank
  4. Return results

Hybrid query

If query includes both text and image:

  • embed both
  • either average/fuse embeddings
  • or retrieve separately and merge scores

Model choice

If you want one model for both text and images

Use a CLIP-like model.

Benefits:

  • simple
  • works well for cross-modal retrieval
  • one embedding space
  • easy operationally

Tradeoff:

  • may be less optimal than specialized models for some domains

If your domain is specialized

Consider:

  • domain-tuned CLIP variants
  • fine-tuning on your own image-text pairs
  • OCR-aware models if documents/screenshots matter

Examples of useful data:

  • image + caption
  • product image + title + description
  • screenshot + extracted text
  • video frame + transcript segment

What to store in the index

For each item, store:

  • id
  • modality = image/text/both
  • embedding or embeddings
  • text fields for BM25 / filters
  • image URL or blob pointer
  • source
  • tags
  • timestamp
  • language
  • permissions / tenant info

If using multiple embeddings per item, common patterns are:

  • one embedding for title
  • one for body text
  • one for image
  • one fused “item” embedding

Important design decisions

1) Single vector vs multiple vectors per item

Single vector

  • easier
  • cheaper
  • less control

Multiple vectors

  • better recall
  • handles long documents and multiple images better
  • more storage and complexity

A good default:

  • one vector for each important field
  • retrieve top-k per field
  • merge results

2) Text chunking

For long text, don’t embed the whole thing as one vector.

Instead:

  • split into chunks of ~200–500 tokens
  • embed each chunk
  • store chunk-level vectors
  • optionally also store a document-level summary embedding

3) OCR and captions matter

If you search screenshots, receipts, slides, or PDFs:

  • run OCR
  • embed OCR text too
  • optionally combine OCR + caption + surrounding metadata

This often boosts retrieval more than image embeddings alone.


4) Add lexical search

Vector search is strong semantically, but lexical search is still valuable.

Best setup:

  • vector search for semantic recall
  • BM25 / keyword index for exact matches
  • hybrid ranking to combine them

This is especially helpful for:

  • names
  • SKUs
  • technical terms
  • product IDs
  • exact phrases

A practical default setup

If you want a simple but strong baseline:

  • Model: CLIP/SigLIP-style shared embedding model
  • Index: vector DB supporting ANN + metadata filters
  • Text augmentation: OCR, captions, titles, tags
  • Hybrid search: vector + BM25
  • Reranking: optional cross-encoder for top 50–200 candidates
  • Chunking: chunk long text into passages
  • Fusing: if item has both text and image, store both and optionally a fused embedding

Example retrieval strategy

For an item with text and image

Store:

  • title_embedding
  • body_chunk_embeddings
  • image_embedding

At query time:

  1. Embed query
  2. Search all fields
  3. Combine scores:
    • title match weighted highest
    • image match moderate
    • body match moderate
  4. Rerank top results

If you’re early-stage

Start with:

  • one shared multimodal embedding model
  • text + image embeddings in same space
  • hybrid BM25 + vector
  • metadata filters
  • reranking later

That gives you a fast path to a working product without overengineering.


Rule of thumb

  • Need easy cross-modal search? Use a shared embedding model.
  • Need best relevance? Use shared embeddings + reranker.
  • Need domain precision? Fine-tune on your own pairs and add OCR/metadata.
  • Need exact matches too? Add BM25.

If you want, I can also sketch:

  1. a reference architecture,
  2. a schema for storing embeddings, or
  3. a model shortlist for your specific use case.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.