Prompt

How do I chunk long documents for better embedding retrieval?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Chunking long documents well is mostly about preserving meaning while keeping chunks small enough to retrieve precisely.

Good chunking principles

1. Chunk by structure first, size second

Use natural boundaries when possible:

  • headings
  • paragraphs
  • sections
  • bullet lists
  • code blocks
  • sentences as a fallback

Avoid splitting in the middle of:

  • a sentence
  • a table row
  • a code block
  • a definition or list item

2. Use overlapping chunks

A small overlap helps preserve context across boundaries.

Typical starting point:

  • chunk size: 200–500 tokens
  • overlap: 10–20% of chunk size

Example:

  • 300-token chunks
  • 50-token overlap

This helps when relevant information is split between adjacent chunks.

3. Match chunk size to the type of content

Different content types benefit from different sizes:

  • FAQ / short answers: 100–200 tokens
  • General prose: 200–500 tokens
  • Technical docs: 300–800 tokens
  • Legal / policy docs: often larger chunks with strong structure
  • Code: chunk by function/class, not arbitrary length

If chunks are too small, you lose context.
If too large, embeddings become diluted and retrieval gets less precise.

4. Keep chunks semantically coherent

Each chunk should ideally answer:

  • “What is this chunk about?”
  • “Would this chunk make sense if retrieved alone?”

If a chunk covers too many topics, split it.
If a topic is spread over several tiny chunks, merge them.

5. Add metadata

Store useful metadata with each chunk:

  • document title
  • section heading
  • subsection heading
  • page number
  • source URL
  • chunk index

This improves retrieval, filtering, and citation quality.

6. Consider hierarchical chunking

For long documents, use multiple levels:

  • small chunks for retrieval
  • larger parent sections for context expansion

This is often called parent-child chunking:

  • retrieve small chunks
  • then fetch the surrounding larger section

It works very well for manuals and reports.


Practical starting recipe

A strong default setup is:

  • split by headings first
  • then by paragraphs
  • then by sentences if needed
  • chunk size: 300–400 tokens
  • overlap: 50–100 tokens
  • keep metadata for section titles and source location

Then test and tune based on retrieval quality.


How to know if your chunks are good

Check retrieval results for:

  • Does the top result contain the answer?
  • Are retrieved chunks too broad?
  • Are important facts split across multiple chunks?
  • Do queries about specific details retrieve the right section?

If answers are incomplete, increase overlap or chunk size slightly.
If irrelevant chunks are returned, reduce chunk size or make chunks more semantically focused.


Common mistakes

  • chunking by fixed character count only
  • using no overlap
  • splitting tables/code awkwardly
  • making chunks too large
  • ignoring document structure
  • not storing metadata

Rule of thumb

If you’re unsure, start with:

  • 400 tokens
  • 50-token overlap
  • structure-aware splitting
  • metadata included

Then evaluate on real queries and adjust.

If you want, I can also give you:

  1. a Python example using LangChain or LlamaIndex, or
  2. a chunking strategy tailored to PDFs, code, or legal docs.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.