Prompt
I’m building a search experience over PDFs and want the best embeddings approach
Latest observation
For PDFs, the “best” embeddings approach is usually not just about the embedding model — it’s about how you split, clean, and retrieve PDF text. A strong default stack looks like this:
Recommended approach
1) Extract text carefully
PDFs are messy. Before embedding:
- remove headers/footers if possible
- fix hyphenation across line breaks
- preserve tables separately if they matter
- keep page and section metadata
- OCR scanned PDFs first if needed
2) Chunk by structure, not just length
Use chunks based on:
- headings / sections / paragraphs
- semantic boundaries
- token budget around 200–500 tokens per chunk, with overlap if needed
For PDFs, structure-aware chunking usually beats fixed-size chunking.
3) Use a strong modern embedding model
Good choices depend on your stack:
Open-source / self-hosted
- bge-large-en-v1.5 — strong general-purpose retrieval
- e5-large-v2 — excellent retrieval quality
- bge-m3 — good if you need multilingual + hybrid use cases
- jina-embeddings-v3 — strong and flexible
Managed APIs
- OpenAI embeddings are a strong default for quality and simplicity
- Cohere embeddings are also very good for retrieval
If you want the highest quality for search, I’d generally prefer:
- bge-m3 if multilingual / hybrid retrieval matters
- e5-large-v2 or bge-large-en-v1.5 for strong English retrieval
- a top-tier API embedding if you want easiest ops and strong out-of-the-box performance
4) Use hybrid retrieval
For PDFs, pure vector search often misses exact terms like:
- product names
- acronyms
- legal clauses
- figure/table references
Best practice:
- combine BM25 / keyword search + vector search
- merge results, then rerank
5) Add a reranker
A reranker often gives the biggest quality boost after chunking. Use a cross-encoder reranker to reorder top results from hybrid retrieval.
6) Query-time enhancements
Helpful techniques:
- query expansion / rewriting
- metadata filters by document, page, date, section
- reranking top 20–100 candidates
- answer citations to page numbers
Best practical default setup
If you want a solid production baseline:
- PDF parsing: structured extraction + OCR if needed
- Chunking: 300–500 tokens, section-aware, 10–20% overlap
- Embeddings:
bge-large-en-v1.5ore5-large-v2 - Retrieval: hybrid BM25 + vector
- Reranking: cross-encoder reranker
- Metadata: store page number, section, document title, table/figure flags
If your PDFs are tricky
Scanned PDFs
- OCR first
- use layout-aware OCR if tables/forms matter
Legal / technical PDFs
- hybrid search is especially important
- chunk by clause/heading
- preserve exact text for citations
Academic papers
- keep abstract, sections, references separate
- use section-aware chunking
Tables-heavy PDFs
- extract tables into structured text or CSV
- embed table summaries separately
My short recommendation
If you want the best overall results for PDF search:
- Use structure-aware chunking
- Use a strong embedding model like bge-large-en-v1.5 / e5-large-v2 / bge-m3
- Combine vector + BM25
- Add reranking
- Preserve page and section metadata for citations
If you want, I can suggest:
- the best embedding model for your exact use case
- a reference architecture
- or a Python implementation for PDF ingestion + search.