Prompt

How do I chunk long PDFs for retrieval-augmented generation?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

For RAG, the goal is to chunk PDFs so each chunk is:

  • small enough to retrieve precisely
  • large enough to preserve meaning
  • structurally clean so the model can cite useful context

Here’s a practical approach.

1) Prefer structure-aware chunking over fixed-size chunks

PDFs often have headings, paragraphs, tables, and page breaks. If possible, chunk by:

  1. Document sections / headings
  2. Paragraphs
  3. Sentences as a fallback

Avoid splitting in the middle of a table row, bullet list, or paragraph unless you must.

2) Use a token-based chunk size

A good starting point:

  • Chunk size: 300–800 tokens
  • Overlap: 10–20% of chunk size
    (e.g. 50–150 tokens overlap)

Typical defaults:

  • 500 tokens with 100-token overlap is a strong baseline
  • If the PDF is highly technical, use smaller chunks like 250–400 tokens
  • If it’s narrative/legal text, 600–1000 tokens may work better

Why token-based? Because embedding models and LLMs operate on tokens, not characters.

3) Add overlap to preserve continuity

Overlap helps when important context straddles boundaries.

Example:

  • Chunk 1: tokens 1–500
  • Chunk 2: tokens 401–900

This improves retrieval for:

  • definitions
  • references like “this method”
  • tables continued across pages
  • lists and procedures

Don’t overdo overlap, or you’ll create redundant embeddings and increase index size.

4) Use hierarchical chunking if the document is long

For large PDFs, use a two-level approach:

  • Parent chunks: larger sections (e.g. 1,500–3,000 tokens)
  • Child chunks: smaller retrievable pieces (e.g. 300–500 tokens)

Retrieve child chunks, but keep the parent section available for expanded context. This is often called parent-child retrieval.

5) Preserve metadata

Store metadata with each chunk:

  • document title
  • page numbers
  • section heading
  • chunk index
  • source filename
  • table/figure labels if relevant

This helps with:

  • citations
  • traceability
  • filtering
  • reranking

Example metadata:

{
  "doc_id": "manual_2024.pdf",
  "page_start": 12,
  "page_end": 13,
  "section": "3.2 Installation",
  "chunk_id": 7
}

6) Handle PDFs carefully during extraction

PDF text extraction can be messy. Clean up:

  • repeated headers/footers
  • page numbers
  • hyphenated line breaks
  • broken columns
  • OCR errors from scanned PDFs

If the PDF is scanned, use OCR first, then chunk the extracted text.

7) Treat tables and bullet lists specially

For:

  • tables: keep row/column structure intact if possible
  • lists: chunk at list boundaries when possible
  • forms / manuals: preserve labels and values together

Sometimes a table should be represented as:

  • a text summary chunk
  • plus a raw table chunk

8) Consider semantic chunking for better retrieval

Instead of splitting by length alone, split when the topic changes. This can improve relevance for retrieval.

Methods:

  • heading-based splitting
  • paragraph similarity
  • embedding-based breakpoint detection

Semantic chunking often works better than naive fixed windows, especially for mixed-format PDFs.

9) Tune chunk size based on your task

Use smaller chunks if users ask:

  • precise factual questions
  • definitions
  • specific clauses

Use larger chunks if users ask:

  • summarization
  • multi-step explanations
  • policy interpretation

10) Evaluate with real queries

Chunking is not one-size-fits-all. Test with:

  • recall@k
  • answer accuracy
  • citation quality
  • chunk relevance

If the retriever misses answers, increase overlap or chunk size.
If it retrieves too much irrelevant text, reduce chunk size or use better semantic splitting.


Good starting recipe

If you want a simple, solid default:

  • Extract text from PDF
  • Remove headers/footers
  • Split by headings/paragraphs
  • Use 500-token chunks
  • Use 100-token overlap
  • Store page and section metadata
  • Use reranking on top of embedding retrieval

Rule of thumb

  • Too small → retrieval is precise but loses context
  • Too large → retrieval returns noisy chunks and embeddings become less specific

A balanced middle ground usually works best.

If you want, I can also give you:

  1. a Python example for PDF chunking, or
  2. a recommended chunking pipeline for LangChain / LlamaIndex / plain OpenAI embeddings.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.