Prompt
How do I chunk documents before generating embeddings for search?
Latest observation
Chunking documents well is mostly about balancing retrieval quality, context limits, and embedding usefulness.
Practical rules of thumb
1) Chunk by semantic boundaries first
Prefer splitting on:
- headings / sections
- paragraphs
- bullet lists
- sentences
Avoid splitting purely by fixed character count if you can. A chunk that ends in the middle of a thought is usually worse for search.
2) Keep chunks moderately sized
A common starting point:
- 200–500 tokens per chunk for general search
- smaller if text is dense and fact-heavy
- larger if meaning depends on more surrounding context
If you’re unsure, start around 300 tokens.
3) Use overlap
Add a small overlap between adjacent chunks so important context at boundaries isn’t lost.
Typical overlap:
- 10–20% of chunk size
- or about 30–80 tokens
Example:
- chunk size: 300 tokens
- overlap: 50 tokens
4) Preserve metadata
Store alongside each chunk:
- document ID
- section/title path
- page number
- chunk index
- source URL / file name
This helps with ranking, filtering, and citations later.
5) Don’t over-chunk
Too many tiny chunks can hurt retrieval:
- embeddings become too specific
- search returns fragments without enough context
- indexing cost increases
6) Don’t under-chunk
Too-large chunks can also hurt:
- embeddings blur multiple topics
- retrieval becomes less precise
A good default strategy
If you want a simple starting point:
- Split by headings/paragraphs
- Merge until each chunk is about 250–400 tokens
- Add 50-token overlap
- Keep chunk boundaries at sentence ends when possible
- Store metadata for each chunk
Example
Original document section:
“Authentication is handled with API keys. Keys should be stored securely in environment variables. Rotate keys every 90 days. If a key is compromised, revoke it immediately.”
Possible chunks:
-
Chunk 1
- Authentication is handled with API keys. Keys should be stored securely in environment variables.
-
Chunk 2
- Keys should be stored securely in environment variables. Rotate keys every 90 days. If a key is compromised, revoke it immediately.
Here the overlap keeps the “stored securely” idea available in both chunks.
Advanced improvements
Hierarchical chunking
Store:
- small chunks for retrieval
- larger parent sections for answer generation
This is useful when you want precise search but fuller context for the final LLM response.
Hybrid chunking
Use:
- structure-aware splitting for docs with headings
- sentence-based splitting for prose
- code-aware splitting for technical docs
Query-aware reranking
Even with good chunking, retrieval improves a lot if you:
- retrieve top 20–50 chunks
- rerank them with a cross-encoder or LLM
- pass only the best few to the generator
What to optimize for
The best chunk size depends on your content:
- FAQ/docs/manuals: smaller chunks work well
- Legal/technical docs: slightly larger chunks with strong overlap
- Narrative text: sentence/paragraph-aware chunks
- Code: split by function/class/module, not arbitrary tokens
If you want, I can also give you:
- a Python chunking function,
- recommended chunk sizes by document type, or
- a LangChain / LlamaIndex setup.
Brands