Prompt
How do I chunk long PDFs for retrieval-augmented generation?
Latest observation
For RAG, the goal is to chunk PDFs so each chunk is:
- small enough to retrieve precisely
- large enough to preserve meaning
- structurally clean so the model can cite useful context
Here’s a practical approach.
1) Prefer structure-aware chunking over fixed-size chunks
PDFs often have headings, paragraphs, tables, and page breaks. If possible, chunk by:
- Document sections / headings
- Paragraphs
- Sentences as a fallback
Avoid splitting in the middle of a table row, bullet list, or paragraph unless you must.
2) Use a token-based chunk size
A good starting point:
- Chunk size: 300–800 tokens
- Overlap: 10–20% of chunk size
(e.g. 50–150 tokens overlap)
Typical defaults:
- 500 tokens with 100-token overlap is a strong baseline
- If the PDF is highly technical, use smaller chunks like 250–400 tokens
- If it’s narrative/legal text, 600–1000 tokens may work better
Why token-based? Because embedding models and LLMs operate on tokens, not characters.
3) Add overlap to preserve continuity
Overlap helps when important context straddles boundaries.
Example:
- Chunk 1: tokens 1–500
- Chunk 2: tokens 401–900
This improves retrieval for:
- definitions
- references like “this method”
- tables continued across pages
- lists and procedures
Don’t overdo overlap, or you’ll create redundant embeddings and increase index size.
4) Use hierarchical chunking if the document is long
For large PDFs, use a two-level approach:
- Parent chunks: larger sections (e.g. 1,500–3,000 tokens)
- Child chunks: smaller retrievable pieces (e.g. 300–500 tokens)
Retrieve child chunks, but keep the parent section available for expanded context. This is often called parent-child retrieval.
5) Preserve metadata
Store metadata with each chunk:
- document title
- page numbers
- section heading
- chunk index
- source filename
- table/figure labels if relevant
This helps with:
- citations
- traceability
- filtering
- reranking
Example metadata:
{
"doc_id": "manual_2024.pdf",
"page_start": 12,
"page_end": 13,
"section": "3.2 Installation",
"chunk_id": 7
}
6) Handle PDFs carefully during extraction
PDF text extraction can be messy. Clean up:
- repeated headers/footers
- page numbers
- hyphenated line breaks
- broken columns
- OCR errors from scanned PDFs
If the PDF is scanned, use OCR first, then chunk the extracted text.
7) Treat tables and bullet lists specially
For:
- tables: keep row/column structure intact if possible
- lists: chunk at list boundaries when possible
- forms / manuals: preserve labels and values together
Sometimes a table should be represented as:
- a text summary chunk
- plus a raw table chunk
8) Consider semantic chunking for better retrieval
Instead of splitting by length alone, split when the topic changes. This can improve relevance for retrieval.
Methods:
- heading-based splitting
- paragraph similarity
- embedding-based breakpoint detection
Semantic chunking often works better than naive fixed windows, especially for mixed-format PDFs.
9) Tune chunk size based on your task
Use smaller chunks if users ask:
- precise factual questions
- definitions
- specific clauses
Use larger chunks if users ask:
- summarization
- multi-step explanations
- policy interpretation
10) Evaluate with real queries
Chunking is not one-size-fits-all. Test with:
- recall@k
- answer accuracy
- citation quality
- chunk relevance
If the retriever misses answers, increase overlap or chunk size.
If it retrieves too much irrelevant text, reduce chunk size or use better semantic splitting.
Good starting recipe
If you want a simple, solid default:
- Extract text from PDF
- Remove headers/footers
- Split by headings/paragraphs
- Use 500-token chunks
- Use 100-token overlap
- Store page and section metadata
- Use reranking on top of embedding retrieval
Rule of thumb
- Too small → retrieval is precise but loses context
- Too large → retrieval returns noisy chunks and embeddings become less specific
A balanced middle ground usually works best.
If you want, I can also give you:
- a Python example for PDF chunking, or
- a recommended chunking pipeline for LangChain / LlamaIndex / plain OpenAI embeddings.
Brands