Prompt
best embedding model for pdf search and retrieval
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For PDF search and retrieval, the “best” embedding model depends on your priorities, but these are the strongest practical choices:
Best overall
OpenAI text-embedding-3-large
- Very strong retrieval quality
- Works well for semantic search over documents
- Good general default for PDFs
- Especially solid if you want a managed API and high accuracy
Best cost/performance
OpenAI text-embedding-3-small
- Cheaper and faster
- Still very good for most PDF search tasks
- Often the best choice if you have lots of documents and want lower cost
Best open-source options
BAAI bge-large-en-v1.5 or bge-m3
bge-large-en-v1.5: strong English retrievalbge-m3: better if you need multilingual or mixed retrieval use cases- Good if you want to self-host
nomic-embed-text-v1.5
- Also a strong open-source option
- Good quality and easy to use
If PDFs are technical / enterprise / multilingual
bge-m3
- Handles multilingual content well
- Useful for mixed-language PDFs and heterogeneous corpora
My practical recommendation
If you want the simplest answer:
- Use
text-embedding-3-largefor best quality - Use
text-embedding-3-smallif cost matters - Use
bge-m3if you need open-source + multilingual
Important note for PDFs
Embedding model quality matters, but PDF retrieval usually depends just as much on:
- chunking strategy
- OCR quality for scanned PDFs
- metadata extraction
- hybrid search (BM25 + embeddings)
- reranking after retrieval
For best results, use:
- good PDF text extraction
- chunking by semantic sections
- embedding model
- reranker
If you want, I can also give you:
- the best model for your budget
- a top 5 comparison table
- or a recommended PDF search stack for production