Prompt
how to chunk internal docs for RAG
Latest observation
For internal-doc RAG, the best chunking strategy is usually structure-aware, semantic, and retrieval-driven rather than a fixed-size only approach.
Practical recommendation
1) Chunk by document structure first
Prefer natural boundaries in this order:
- Section / subsection
- Paragraph
- Bullet list / numbered list
- Table row group
- Code block / policy clause / FAQ item
This preserves meaning better than splitting blindly by token count.
2) Then cap chunk size
A good starting point for internal docs:
- Chunk size: 300–800 tokens
- Overlap: 50–150 tokens
Smaller chunks help precision; larger chunks help context.
For policy docs, SOPs, legal, or technical guides, 400–600 tokens is often a strong default.
3) Keep chunks self-contained
Each chunk should ideally include:
- Title / heading
- Parent section heading if needed
- Any definitions or key references used in the chunk
Example:
Document: Benefits Policy
Section: Parental Leave
Subsection: Eligibility
[chunk text...]
4) Use semantic splitting for long sections
If a section is long, split by topic shifts rather than exact length.
Good indicators:
- New bullet group
- Change in subject
- New numbered step
- New paragraph with a different intent
5) Add metadata aggressively
Store metadata with each chunk:
doc_iddoc_titlesection_pathcreated_atowner/teampermissionssource_urlchunk_index
This is critical for filtering and access control in internal RAG.
Best practices by doc type
Policies / HR / Legal
- Chunk by clause or subsection
- Keep exact wording intact
- Avoid splitting definitions from their usage
- Use smaller chunks, since precision matters
Technical docs / runbooks
- Chunk by procedure step or subprocedure
- Keep code blocks intact
- Include prerequisites and warnings with the relevant steps
Meeting notes / wikis
- Chunk by topic or heading
- Remove duplicated boilerplate
- Use semantic chunking if the structure is messy
Product specs
- Chunk by feature, requirement, or acceptance criteria
- Keep tables and requirements together
- Include related context from the parent section
A strong default pipeline
- Parse document structure
- Split on headings
- Within each section, split by paragraphs
- Merge until target token size
- Apply small overlap
- Attach heading metadata
- Index chunks in vector store
- Optionally re-rank retrieved chunks
Retrieval-specific improvements
Hybrid chunking
Use:
- small chunks for exact retrieval
- larger parent chunks for context after retrieval
This is often called parent-child retrieval:
- Embed smaller child chunks
- Retrieve them
- Expand to the parent section when answering
Hierarchical indexing
Store:
- document
- section
- chunk
This helps answer broad and narrow questions.
Use reranking
A reranker can compensate for imperfect chunking and improve top-k results.
Common mistakes
- Making chunks too large
- Splitting in the middle of tables or code
- Ignoring headings
- No metadata
- Using the same chunk size for every doc type
- Not testing chunking with real queries
Rule of thumb
If users ask:
- “What is the policy on X?” → smaller, exact chunks
- “How do I do X?” → procedure-aligned chunks
- “Summarize this doc” → larger chunks or hierarchical retrieval
Recommended starting config
If you want a simple baseline:
- Split by headings first
- Target 500 tokens
- Overlap 100 tokens
- Keep tables/code blocks intact
- Add section-path metadata
- Use parent-child retrieval if docs are long
If you want, I can also give you:
- a Python chunking recipe,
- a chunking strategy by doc type, or
- a LangChain / LlamaIndex implementation.
Brands