Prompt
How do I chunk long documents for better embedding retrieval?
Latest observation
Chunking long documents well is mostly about preserving meaning while keeping chunks small enough to retrieve precisely.
Good chunking principles
1. Chunk by structure first, size second
Use natural boundaries when possible:
- headings
- paragraphs
- sections
- bullet lists
- code blocks
- sentences as a fallback
Avoid splitting in the middle of:
- a sentence
- a table row
- a code block
- a definition or list item
2. Use overlapping chunks
A small overlap helps preserve context across boundaries.
Typical starting point:
- chunk size: 200–500 tokens
- overlap: 10–20% of chunk size
Example:
- 300-token chunks
- 50-token overlap
This helps when relevant information is split between adjacent chunks.
3. Match chunk size to the type of content
Different content types benefit from different sizes:
- FAQ / short answers: 100–200 tokens
- General prose: 200–500 tokens
- Technical docs: 300–800 tokens
- Legal / policy docs: often larger chunks with strong structure
- Code: chunk by function/class, not arbitrary length
If chunks are too small, you lose context.
If too large, embeddings become diluted and retrieval gets less precise.
4. Keep chunks semantically coherent
Each chunk should ideally answer:
- “What is this chunk about?”
- “Would this chunk make sense if retrieved alone?”
If a chunk covers too many topics, split it.
If a topic is spread over several tiny chunks, merge them.
5. Add metadata
Store useful metadata with each chunk:
- document title
- section heading
- subsection heading
- page number
- source URL
- chunk index
This improves retrieval, filtering, and citation quality.
6. Consider hierarchical chunking
For long documents, use multiple levels:
- small chunks for retrieval
- larger parent sections for context expansion
This is often called parent-child chunking:
- retrieve small chunks
- then fetch the surrounding larger section
It works very well for manuals and reports.
Practical starting recipe
A strong default setup is:
- split by headings first
- then by paragraphs
- then by sentences if needed
- chunk size: 300–400 tokens
- overlap: 50–100 tokens
- keep metadata for section titles and source location
Then test and tune based on retrieval quality.
How to know if your chunks are good
Check retrieval results for:
- Does the top result contain the answer?
- Are retrieved chunks too broad?
- Are important facts split across multiple chunks?
- Do queries about specific details retrieve the right section?
If answers are incomplete, increase overlap or chunk size slightly.
If irrelevant chunks are returned, reduce chunk size or make chunks more semantically focused.
Common mistakes
- chunking by fixed character count only
- using no overlap
- splitting tables/code awkwardly
- making chunks too large
- ignoring document structure
- not storing metadata
Rule of thumb
If you’re unsure, start with:
- 400 tokens
- 50-token overlap
- structure-aware splitting
- metadata included
Then evaluate on real queries and adjust.
If you want, I can also give you:
- a Python example using LangChain or LlamaIndex, or
- a chunking strategy tailored to PDFs, code, or legal docs.
Brands