Prompt
What should I use for embeddings for PDF search and document retrieval?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For PDF search / document retrieval, the best embedding model depends on whether you want:
- Simple semantic search
- High-quality retrieval across lots of documents
- Low cost / fast inference
- On-device / open-source deployment
Good default choices
If you want the easiest strong baseline
- OpenAI text-embedding-3-small
- Great quality
- Cheap
- Fast
- Usually a very solid default for PDF search
If you want higher quality and can pay more
- OpenAI text-embedding-3-large
- Better retrieval quality
- Useful when search precision matters more
- Good for complex, heterogeneous documents
Strong open-source options
- BAAI/bge-large-en-v1.5 for English
- BAAI/bge-m3 for multilingual + hybrid retrieval use cases
- intfloat/e5-large-v2 as another excellent general-purpose choice
What I’d recommend in practice
For most applications
Use:
- text-embedding-3-small if you want a low-cost, strong baseline
- text-embedding-3-large if quality is more important than cost
If your PDFs are mostly technical, legal, or long-form docs
Consider:
- bge-m3 or e5-large-v2
- These work well in retrieval pipelines, especially if you chunk documents well
Important: embeddings alone aren’t enough
For PDF search, results depend heavily on:
1. Text extraction quality
- Use a good PDF parser
- OCR scanned PDFs if needed
- Preserve headings, tables, and section structure where possible
2. Chunking strategy
- Split documents into meaningful chunks
- Common chunk size: 200–500 tokens
- Add overlap: ~10–20%
- Keep metadata like:
- page number
- section title
- document name
3. Retrieval method
Best results often come from:
- Dense vector search with embeddings
- plus keyword/BM25 search
- optionally reranking with a cross-encoder or LLM reranker
This is often called hybrid search, and it performs better than embeddings alone for PDFs.
Quick recommendations by scenario
Small project / prototype
- text-embedding-3-small
- Vector DB: FAISS, Pinecone, Weaviate, pgvector
Best quality with OpenAI
- text-embedding-3-large
- Add reranking if possible
Fully open-source
- bge-m3
- Or e5-large-v2 for English
- Use hybrid search if you can
Multilingual PDFs
- bge-m3 is a strong choice
Rule of thumb
If you’re unsure, start with:
text-embedding-3-small + good chunking + metadata + reranking
Then upgrade to text-embedding-3-large if you need better retrieval accuracy.
If you want, I can also recommend:
- the best embedding model for your budget
- a PDF chunking strategy
- or a full RAG stack for PDF search.