Prompt

embedding chunk size for rag

Artificial Intelligence / AI Search2 observationsLast seen Oct 11, 2026

Latest observation

Oct 11, 2026OpenAI APIWeb search: off

For RAG, a good default embedding chunk size is:

  • 300–800 tokens per chunk
  • A very common starting point: ~500 tokens
  • Use 10–20% overlap between chunks

Quick rule of thumb

  • Smaller chunks = better precision, more embeddings, more context fragmentation
  • Larger chunks = more context per chunk, but noisier retrieval and less exact matches

Practical defaults by content type

  • FAQ / short docs: 200–400 tokens
  • Technical docs / manuals: 400–800 tokens
  • Legal / policy / long-form prose: 800–1200 tokens if sections are coherent
  • Code: chunk by function/class, often 100–300 lines or semantic blocks

Important

Chunk by semantic boundaries when possible:

  • headings
  • paragraphs
  • sections
  • function/class boundaries

Avoid splitting in the middle of ideas just to hit a token count.

Best starting configuration

If you just want a solid baseline:

  • Chunk size: 500 tokens
  • Overlap: 50–100 tokens
  • Split on headings/paragraphs first

Tune based on retrieval results

If answers miss needed context:

  • increase chunk size or overlap

If retrieval is too noisy:

  • decrease chunk size

If you want, I can recommend a chunking strategy based on your data type (PDFs, web pages, code, chat logs, etc.).

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (2 observations for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.