Prompt

I'm building a RAG pipeline and need help choosing embeddings and chunking strategy

Artificial Intelligence / AI Search1 observationLast seen Oct 1, 2026

Latest observation

Oct 1, 2026OpenAI APIWeb search: off

Absolutely — for a RAG pipeline, embeddings and chunking matter a lot more than people often expect. The “best” choice depends on your content type, query style, latency budget, and whether you need multilingual support.

Quick recommendation

If you want a strong default:

  • Embeddings:

    • OpenAI text-embedding-3-large for highest quality
    • OpenAI text-embedding-3-small for cheaper/faster and often good enough
    • If self-hosting: bge-large-en-v1.5 or e5-large-v2 for English
    • For multilingual: bge-m3 is a strong option
  • Chunking strategy:

    • Start with semantic-ish chunks around 300–800 tokens
    • Use 10–20% overlap if the text is narrative or contains definitions/procedures
    • Prefer structure-aware chunking: split by headings, paragraphs, bullet lists, code blocks, etc.
    • Avoid arbitrary fixed-size splitting unless the data is very uniform

How to choose embeddings

1) Consider your primary retrieval style

If your queries are short and factual

Examples:

  • “What is the refund policy?”
  • “How do I rotate logs?”
  • “What is the SLA for enterprise customers?”

Use embeddings that do well on semantic similarity and short-query matching:

  • OpenAI text-embedding-3-large
  • bge-large-en-v1.5
  • e5-large-v2

If your content is technical or code-heavy

You may want:

  • A model that handles exact terminology well
  • Potentially hybrid search with BM25 + embeddings

Good options:

  • bge-large-en-v1.5
  • e5-large-v2
  • bge-m3 if multilingual/heterogeneous corpus

If your corpus is multilingual

Use a multilingual model rather than translating everything:

  • bge-m3
  • intfloat/multilingual-e5-large
  • OpenAI embeddings also handle multilingual reasonably well, but dedicated multilingual models are often stronger for retrieval in non-English corpora

2) Tradeoff: quality vs cost vs latency

Best quality

  • text-embedding-3-large
  • bge-large-en-v1.5

Better cost/performance

  • text-embedding-3-small
  • bge-base-en-v1.5
  • e5-base-v2

Self-hosting considerations

If you self-host embeddings:

  • GPU/CPU inference cost
  • Throughput
  • Index rebuild time
  • Dimensionality and vector store cost

3) Use one embedding model for retrieval, not multiple unless needed

Don’t overcomplicate at first. Pick one strong retriever embedding model and evaluate.

A common stack:

  • Dense embeddings for semantic retrieval
  • BM25 keyword retrieval
  • Optional reranker on top

This is often better than trying to use multiple embedding models at once.


Chunking strategy: what works best

1) Chunk by structure first

Best practice is to split along:

  • Markdown headings
  • HTML sections
  • PDF paragraphs
  • Document sections/subsections
  • Bullet lists
  • Code blocks separately

Why this helps:

  • Keeps related information together
  • Makes retrieved context more coherent
  • Reduces “orphaned” fragments

2) Recommended chunk sizes

A good starting range:

  • Narrative/docs/help content: 300–600 tokens
  • Technical docs / policies / manuals: 400–800 tokens
  • FAQ-style content: often one Q/A per chunk, or slightly larger if answer spans multiple paragraphs
  • Code/docs: chunk by function/class/module boundaries when possible

Overlap

  • 10–20% overlap is usually enough
  • Example: 500-token chunks with 50–100 token overlap

Too much overlap:

  • Increases index size
  • Creates near-duplicate chunks
  • Can hurt diversity in retrieval results

Too little overlap:

  • Risks splitting key context across chunks

3) Avoid purely fixed-size chunking unless necessary

Fixed-size chunking can split:

  • A definition from its explanation
  • A question from its answer
  • A procedure step from its exception
  • A code function from its docstring

Use fixed-size chunking only as a fallback.


4) Parent-child chunking is often excellent

This is a strong RAG pattern:

  • Child chunks: small chunks embedded and retrieved
  • Parent chunks: larger surrounding sections returned to the LLM

Benefits:

  • High recall from small chunks
  • Good context from larger chunks

Example:

  • Embed 200–400 token chunks
  • Store a mapping to 1,000–2,000 token parent sections
  • Retrieve small chunks, then expand to parent section for generation

This works especially well for:

  • Policies
  • Technical documentation
  • Legal content
  • Long articles

Best chunking approach by content type

Documentation / knowledge base

Use:

  • Heading-based splitting
  • Paragraph-based splitting
  • 400–800 token chunks
  • 10–15% overlap

FAQs

Use:

  • One question-answer per chunk
  • Keep each answer intact if possible
  • Add metadata with the question/topic

PDFs

PDFs are tricky because layout can be messy. Use:

  • Layout-aware parsing
  • Section/paragraph reconstruction
  • Remove headers/footers
  • Chunk by logical sections, not page breaks

Code repositories

Use:

  • Function/class-level chunking
  • Include surrounding docstrings/comments
  • Possibly store file path, symbol name, and language metadata

Legal/policy docs

Use:

  • Clause/section-based chunking
  • Preserve numbering and headings
  • Prefer smaller chunks with parent expansion

Retrieval quality tips beyond embeddings and chunking

1) Add metadata

Store metadata like:

  • document title
  • section heading
  • URL/path
  • timestamp/version
  • language
  • source type

This improves filtering and answer grounding.

2) Use hybrid retrieval

Dense embeddings are great, but keyword search helps with:

  • exact names
  • product IDs
  • error codes
  • legal terms
  • acronyms

A strong baseline:

  • BM25 + dense vectors
  • Then rerank top results

3) Add a reranker if you can

A reranker can dramatically improve top-k quality.

Pipeline:

  1. Retrieve top 20–50 with embeddings/BM25
  2. Rerank with a cross-encoder or LLM-based reranker
  3. Pass top 3–8 chunks to the generation model

This often matters more than squeezing the last bit of performance from embeddings.


Practical starting configs

Option A: Strong general-purpose RAG

  • Embeddings: text-embedding-3-large
  • Chunking: 500 tokens, 75 token overlap
  • Split on headings/paragraphs first
  • Hybrid retrieval + reranker

Option B: Lower-cost production setup

  • Embeddings: text-embedding-3-small
  • Chunking: 400–600 tokens, 50–75 token overlap
  • Hybrid retrieval
  • Reranker if budget allows

Option C: Self-hosted open-source English

  • Embeddings: bge-large-en-v1.5 or e5-large-v2
  • Chunking: structure-aware, 400–700 tokens
  • Hybrid retrieval + reranker

Option D: Multilingual

  • Embeddings: bge-m3
  • Chunking: structure-aware, 400–700 tokens
  • Hybrid retrieval
  • Reranker with multilingual support

How to evaluate

Don’t choose blindly. Test on a small labeled set.

Measure:

  • Recall@k: does the correct chunk appear in top k?
  • MRR / nDCG: how high is the right chunk ranked?
  • Answer accuracy: does final response actually answer correctly?
  • Citation precision: are retrieved chunks relevant and well-grounded?

Create 20–100 real queries and inspect:

  • retrieved chunks
  • whether the answer is contained
  • whether chunk boundaries are helping or hurting

A simple rule of thumb

If you’re unsure:

  • Use structure-aware chunks
  • Start around 500 tokens
  • Overlap ~75 tokens
  • Use text-embedding-3-large or bge-large-en-v1.5
  • Add hybrid retrieval
  • Add a reranker

That’s a very solid baseline.

If you want, I can also help you choose based on your exact setup:

  • document type
  • language(s)
  • expected query types
  • vector DB
  • latency/cost constraints

If you share those, I can recommend a concrete embedding model + chunking configuration.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.