Prompt

How do I chunk documents for embeddings so search returns the right passages?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To make search return the right passages, chunking should optimize for retrieval granularity, not just readability.

Practical chunking rules

1) Chunk by semantic boundaries first

Prefer splitting at:

  • section headings
  • paragraphs
  • bullet groups
  • code blocks / tables as whole units

Avoid splitting in the middle of:

  • a sentence
  • a list item
  • a table row
  • a definition + its explanation

2) Keep chunks small enough to be specific

A good starting point is:

  • 150–400 tokens per chunk for prose
  • larger for dense technical text if needed
  • overlap of 10–20% between chunks

Why: if chunks are too large, embeddings blur multiple topics and search returns broad, noisy passages.

3) Add overlap, but not too much

Use overlap to preserve context across boundaries:

  • 20–50 tokens is often enough
  • more overlap can create duplicate retrieval results

Example:

  • Chunk 1: tokens 1–250
  • Chunk 2: tokens 220–470
  • Chunk 3: tokens 440–690

4) Preserve metadata with each chunk

Store:

  • document title
  • section heading
  • page number
  • chunk index
  • source URL / file path

This helps ranking and lets you reconstruct context after retrieval.

5) Don’t mix unrelated content in one chunk

Bad:

  • a heading from one topic
  • followed by a paragraph from another
  • plus a footer or sidebar text

Good:

  • one coherent section per chunk

6) Use hierarchical chunking for long docs

For long documents:

  • create small chunks for embeddings/retrieval
  • also store parent sections for context

This is often called parent-child retrieval:

  • retrieve a small chunk
  • return the surrounding larger section if needed

This improves precision without losing context.


A strong default strategy

If you want a simple starting point:

  • Split on headings/paragraphs
  • Target 250 tokens
  • Use 50-token overlap
  • Keep code blocks, tables, and list items intact
  • Attach metadata for section/page/source

That works well for many docs.


When to change chunk size

Use smaller chunks if:

  • queries are very specific
  • docs have lots of unrelated information
  • search returns passages that are too broad

Use larger chunks if:

  • passages need more surrounding context
  • text is highly compressed or technical
  • important meaning spans multiple paragraphs

Common mistakes

  • Chunking by fixed character count only
  • Using huge chunks “to preserve context”
  • No overlap
  • Dropping headings and metadata
  • Splitting tables, code, or lists mid-way
  • Embedding OCR noise, headers, and footers together with content

Best practice: tune with evaluation

Test your chunking by asking:

  • Does the top retrieved chunk directly answer the query?
  • Is the answer contained in the chunk without too much unrelated text?
  • Do adjacent chunks duplicate too much?

If retrieval is noisy, decrease chunk size. If retrieval misses needed context, increase overlap or use parent-child retrieval.

If you want, I can give you:

  1. a recommended chunking recipe by document type (PDFs, manuals, code, FAQs), or
  2. a Python example for chunking text for embeddings.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.