Prompt

How do I chunk documents before generating embeddings for search?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Chunking documents well is mostly about balancing retrieval quality, context limits, and embedding usefulness.

Practical rules of thumb

1) Chunk by semantic boundaries first

Prefer splitting on:

  • headings / sections
  • paragraphs
  • bullet lists
  • sentences

Avoid splitting purely by fixed character count if you can. A chunk that ends in the middle of a thought is usually worse for search.

2) Keep chunks moderately sized

A common starting point:

  • 200–500 tokens per chunk for general search
  • smaller if text is dense and fact-heavy
  • larger if meaning depends on more surrounding context

If you’re unsure, start around 300 tokens.

3) Use overlap

Add a small overlap between adjacent chunks so important context at boundaries isn’t lost.

Typical overlap:

  • 10–20% of chunk size
  • or about 30–80 tokens

Example:

  • chunk size: 300 tokens
  • overlap: 50 tokens

4) Preserve metadata

Store alongside each chunk:

  • document ID
  • section/title path
  • page number
  • chunk index
  • source URL / file name

This helps with ranking, filtering, and citations later.

5) Don’t over-chunk

Too many tiny chunks can hurt retrieval:

  • embeddings become too specific
  • search returns fragments without enough context
  • indexing cost increases

6) Don’t under-chunk

Too-large chunks can also hurt:

  • embeddings blur multiple topics
  • retrieval becomes less precise

A good default strategy

If you want a simple starting point:

  • Split by headings/paragraphs
  • Merge until each chunk is about 250–400 tokens
  • Add 50-token overlap
  • Keep chunk boundaries at sentence ends when possible
  • Store metadata for each chunk

Example

Original document section:

“Authentication is handled with API keys. Keys should be stored securely in environment variables. Rotate keys every 90 days. If a key is compromised, revoke it immediately.”

Possible chunks:

  1. Chunk 1

    • Authentication is handled with API keys. Keys should be stored securely in environment variables.
  2. Chunk 2

    • Keys should be stored securely in environment variables. Rotate keys every 90 days. If a key is compromised, revoke it immediately.

Here the overlap keeps the “stored securely” idea available in both chunks.


Advanced improvements

Hierarchical chunking

Store:

  • small chunks for retrieval
  • larger parent sections for answer generation

This is useful when you want precise search but fuller context for the final LLM response.

Hybrid chunking

Use:

  • structure-aware splitting for docs with headings
  • sentence-based splitting for prose
  • code-aware splitting for technical docs

Query-aware reranking

Even with good chunking, retrieval improves a lot if you:

  • retrieve top 20–50 chunks
  • rerank them with a cross-encoder or LLM
  • pass only the best few to the generator

What to optimize for

The best chunk size depends on your content:

  • FAQ/docs/manuals: smaller chunks work well
  • Legal/technical docs: slightly larger chunks with strong overlap
  • Narrative text: sentence/paragraph-aware chunks
  • Code: split by function/class/module, not arbitrary tokens

If you want, I can also give you:

  1. a Python chunking function,
  2. recommended chunk sizes by document type, or
  3. a LangChain / LlamaIndex setup.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.