Prompt
How do I chunk documents for embeddings so search returns the right passages?
Latest observation
To make search return the right passages, chunking should optimize for retrieval granularity, not just readability.
Practical chunking rules
1) Chunk by semantic boundaries first
Prefer splitting at:
- section headings
- paragraphs
- bullet groups
- code blocks / tables as whole units
Avoid splitting in the middle of:
- a sentence
- a list item
- a table row
- a definition + its explanation
2) Keep chunks small enough to be specific
A good starting point is:
- 150–400 tokens per chunk for prose
- larger for dense technical text if needed
- overlap of 10–20% between chunks
Why: if chunks are too large, embeddings blur multiple topics and search returns broad, noisy passages.
3) Add overlap, but not too much
Use overlap to preserve context across boundaries:
- 20–50 tokens is often enough
- more overlap can create duplicate retrieval results
Example:
- Chunk 1: tokens 1–250
- Chunk 2: tokens 220–470
- Chunk 3: tokens 440–690
4) Preserve metadata with each chunk
Store:
- document title
- section heading
- page number
- chunk index
- source URL / file path
This helps ranking and lets you reconstruct context after retrieval.
5) Don’t mix unrelated content in one chunk
Bad:
- a heading from one topic
- followed by a paragraph from another
- plus a footer or sidebar text
Good:
- one coherent section per chunk
6) Use hierarchical chunking for long docs
For long documents:
- create small chunks for embeddings/retrieval
- also store parent sections for context
This is often called parent-child retrieval:
- retrieve a small chunk
- return the surrounding larger section if needed
This improves precision without losing context.
A strong default strategy
If you want a simple starting point:
- Split on headings/paragraphs
- Target 250 tokens
- Use 50-token overlap
- Keep code blocks, tables, and list items intact
- Attach metadata for section/page/source
That works well for many docs.
When to change chunk size
Use smaller chunks if:
- queries are very specific
- docs have lots of unrelated information
- search returns passages that are too broad
Use larger chunks if:
- passages need more surrounding context
- text is highly compressed or technical
- important meaning spans multiple paragraphs
Common mistakes
- Chunking by fixed character count only
- Using huge chunks “to preserve context”
- No overlap
- Dropping headings and metadata
- Splitting tables, code, or lists mid-way
- Embedding OCR noise, headers, and footers together with content
Best practice: tune with evaluation
Test your chunking by asking:
- Does the top retrieved chunk directly answer the query?
- Is the answer contained in the chunk without too much unrelated text?
- Do adjacent chunks duplicate too much?
If retrieval is noisy, decrease chunk size. If retrieval misses needed context, increase overlap or use parent-child retrieval.
If you want, I can give you:
- a recommended chunking recipe by document type (PDFs, manuals, code, FAQs), or
- a Python example for chunking text for embeddings.