Prompt

I’m building a search experience over PDFs and want the best embeddings approach

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

For PDFs, the “best” embeddings approach is usually not just about the embedding model — it’s about how you split, clean, and retrieve PDF text. A strong default stack looks like this:

Recommended approach

1) Extract text carefully

PDFs are messy. Before embedding:

  • remove headers/footers if possible
  • fix hyphenation across line breaks
  • preserve tables separately if they matter
  • keep page and section metadata
  • OCR scanned PDFs first if needed

2) Chunk by structure, not just length

Use chunks based on:

  • headings / sections / paragraphs
  • semantic boundaries
  • token budget around 200–500 tokens per chunk, with overlap if needed

For PDFs, structure-aware chunking usually beats fixed-size chunking.

3) Use a strong modern embedding model

Good choices depend on your stack:

Open-source / self-hosted

  • bge-large-en-v1.5 — strong general-purpose retrieval
  • e5-large-v2 — excellent retrieval quality
  • bge-m3 — good if you need multilingual + hybrid use cases
  • jina-embeddings-v3 — strong and flexible

Managed APIs

  • OpenAI embeddings are a strong default for quality and simplicity
  • Cohere embeddings are also very good for retrieval

If you want the highest quality for search, I’d generally prefer:

  • bge-m3 if multilingual / hybrid retrieval matters
  • e5-large-v2 or bge-large-en-v1.5 for strong English retrieval
  • a top-tier API embedding if you want easiest ops and strong out-of-the-box performance

4) Use hybrid retrieval

For PDFs, pure vector search often misses exact terms like:

  • product names
  • acronyms
  • legal clauses
  • figure/table references

Best practice:

  • combine BM25 / keyword search + vector search
  • merge results, then rerank

5) Add a reranker

A reranker often gives the biggest quality boost after chunking. Use a cross-encoder reranker to reorder top results from hybrid retrieval.

6) Query-time enhancements

Helpful techniques:

  • query expansion / rewriting
  • metadata filters by document, page, date, section
  • reranking top 20–100 candidates
  • answer citations to page numbers

Best practical default setup

If you want a solid production baseline:

  • PDF parsing: structured extraction + OCR if needed
  • Chunking: 300–500 tokens, section-aware, 10–20% overlap
  • Embeddings: bge-large-en-v1.5 or e5-large-v2
  • Retrieval: hybrid BM25 + vector
  • Reranking: cross-encoder reranker
  • Metadata: store page number, section, document title, table/figure flags

If your PDFs are tricky

Scanned PDFs

  • OCR first
  • use layout-aware OCR if tables/forms matter

Legal / technical PDFs

  • hybrid search is especially important
  • chunk by clause/heading
  • preserve exact text for citations

Academic papers

  • keep abstract, sections, references separate
  • use section-aware chunking

Tables-heavy PDFs

  • extract tables into structured text or CSV
  • embed table summaries separately

My short recommendation

If you want the best overall results for PDF search:

  1. Use structure-aware chunking
  2. Use a strong embedding model like bge-large-en-v1.5 / e5-large-v2 / bge-m3
  3. Combine vector + BM25
  4. Add reranking
  5. Preserve page and section metadata for citations

If you want, I can suggest:

  • the best embedding model for your exact use case
  • a reference architecture
  • or a Python implementation for PDF ingestion + search.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.