Prompt
what should I use to search across PDFs and images semantically?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If you want to search across PDFs and images semantically (not just keyword matching), you usually want a multimodal retrieval stack:
Best approach
-
Extract text from PDFs
- If PDFs have selectable text: use normal text extraction.
- If they’re scanned: run OCR first.
-
Process images with OCR and/or vision embeddings
- For text in images: OCR.
- For general visual similarity or content search: use image embeddings.
-
Store everything in a vector database
- Chunk PDF text and OCR output.
- Generate embeddings for both text and images.
- Put them in a vector DB for semantic search.
Good model choices
For text/PDF semantic search
- OpenAI text embeddings: strong general-purpose embeddings
- BGE / E5: popular open-source alternatives
- Cohere Embed: also solid
For images
- CLIP / SigLIP: classic image-text embedding models
- OpenAI multimodal embeddings if you want a managed API and unified search
- BLIP / Florence / other vision-language models for richer understanding, depending on your use case
For OCR
- Tesseract: free, but weaker on messy docs
- PaddleOCR: often better for real-world scans
- AWS Textract / Google Document AI / Azure OCR: strong managed options
Vector databases
- Pinecone: managed, easy scaling
- Weaviate: good multimodal support
- Milvus: powerful open-source
- Qdrant: simple and popular
- FAISS: good if you want local/index-only
Practical recommendation
If you want the simplest robust setup:
- PDF text: extract text with PyMuPDF / pdfplumber
- Scanned PDFs + images: OCR with PaddleOCR or Textract
- Embeddings:
- text: OpenAI embeddings or BGE
- images: CLIP/SigLIP or a multimodal embedding model
- Search backend: Qdrant or Weaviate
If you want one unified semantic search over both text and images
Look for a shared embedding space:
- CLIP-style models are ideal because both images and text map into the same vector space.
- That makes cross-modal search easier, e.g.:
- search with text, retrieve images
- search with image, retrieve text-heavy docs
Rule of thumb
- If documents are mostly textual: OCR + text embeddings + vector DB
- If you need true cross-modal search: OCR + CLIP-style embeddings + vector DB
- If documents are complex layouts: add document parsing tools like
docling,unstructured, orlayoutparser
If you want, I can recommend a specific stack for:
- local/open-source only
- best managed/cloud
- highest accuracy
- cheapest setup