Prompt
need image and text embeddings in one pipeline
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If you want image + text embeddings in one pipeline, the usual solution is to use a multimodal model that puts both modalities into the same vector space.
Good options
1) CLIP / SigLIP / OpenCLIP
- Text embedding: encode text with the text encoder
- Image embedding: encode image with the image encoder
- Both embeddings are aligned, so you can compare them directly with cosine similarity
This is the most common choice for:
- image-text search
- retrieval
- matching captions to images
- zero-shot classification
2) Multimodal encoders in Hugging Face
You can use models like:
openai/clip-vit-base-patch32laion/CLIP-ViT-B-32-laion2B-s34B-b79Kgoogle/siglip-base-patch16-224
Simple pipeline
- Load one multimodal model
- Preprocess image and text using the model’s processor
- Run inference
- Get:
image_embedstext_embeds
- Normalize vectors if needed
- Use cosine similarity or store embeddings in a vector DB
Example in Python with Hugging Face
from transformers import CLIPProcessor, CLIPModel
from PIL import Image
import torch
model_name = "openai/clip-vit-base-patch32"
model = CLIPModel.from_pretrained(model_name)
processor = CLIPProcessor.from_pretrained(model_name)
image = Image.open("image.jpg")
text = ["a dog playing in the park"]
inputs = processor(text=text, images=image, return_tensors="pt", padding=True)
outputs = model(**inputs)
image_embeds = outputs.image_embeds
text_embeds = outputs.text_embeds
# Normalize
image_embeds = image_embeds / image_embeds.norm(dim=-1, keepdim=True)
text_embeds = text_embeds / text_embeds.norm(dim=-1, keepdim=True)
similarity = (image_embeds @ text_embeds.T).item()
print("Similarity:", similarity)
If you need “one pipeline” for both generation and embeddings
If by “one pipeline” you mean:
- extract embeddings
- and also use them for downstream tasks
then CLIP-style models are still the best fit. They are not generative, but they are very efficient for embedding retrieval workflows.
If you want a single encoder for both text and images
You can also use:
- multimodal transformers
- joint embedding models
- VLMs like BLIP/Florence, but those are usually better for captioning/VQA than fast embedding retrieval
Recommendation
If your goal is shared image + text embeddings, use:
CLIP / OpenCLIP / SigLIP
If you want, I can give you:
- a FastAPI pipeline
- a PyTorch batch embedding pipeline
- a sentence-transformers style implementation
- a vector database ingestion example