Prompt
embedding chunk size for rag
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
For RAG, a good default embedding chunk size is:
- 300–800 tokens per chunk
- A very common starting point: ~500 tokens
- Use 10–20% overlap between chunks
Quick rule of thumb
- Smaller chunks = better precision, more embeddings, more context fragmentation
- Larger chunks = more context per chunk, but noisier retrieval and less exact matches
Practical defaults by content type
- FAQ / short docs: 200–400 tokens
- Technical docs / manuals: 400–800 tokens
- Legal / policy / long-form prose: 800–1200 tokens if sections are coherent
- Code: chunk by function/class, often 100–300 lines or semantic blocks
Important
Chunk by semantic boundaries when possible:
- headings
- paragraphs
- sections
- function/class boundaries
Avoid splitting in the middle of ideas just to hit a token count.
Best starting configuration
If you just want a solid baseline:
- Chunk size: 500 tokens
- Overlap: 50–100 tokens
- Split on headings/paragraphs first
Tune based on retrieval results
If answers miss needed context:
- increase chunk size or overlap
If retrieval is too noisy:
- decrease chunk size
If you want, I can recommend a chunking strategy based on your data type (PDFs, web pages, code, chat logs, etc.).