Prompt
I'm building a RAG pipeline and need help choosing embeddings and chunking strategy
Latest observation
Absolutely — for a RAG pipeline, embeddings and chunking matter a lot more than people often expect. The “best” choice depends on your content type, query style, latency budget, and whether you need multilingual support.
Quick recommendation
If you want a strong default:
-
Embeddings:
- OpenAI
text-embedding-3-largefor highest quality - OpenAI
text-embedding-3-smallfor cheaper/faster and often good enough - If self-hosting:
bge-large-en-v1.5ore5-large-v2for English - For multilingual:
bge-m3is a strong option
- OpenAI
-
Chunking strategy:
- Start with semantic-ish chunks around 300–800 tokens
- Use 10–20% overlap if the text is narrative or contains definitions/procedures
- Prefer structure-aware chunking: split by headings, paragraphs, bullet lists, code blocks, etc.
- Avoid arbitrary fixed-size splitting unless the data is very uniform
How to choose embeddings
1) Consider your primary retrieval style
If your queries are short and factual
Examples:
- “What is the refund policy?”
- “How do I rotate logs?”
- “What is the SLA for enterprise customers?”
Use embeddings that do well on semantic similarity and short-query matching:
- OpenAI
text-embedding-3-large bge-large-en-v1.5e5-large-v2
If your content is technical or code-heavy
You may want:
- A model that handles exact terminology well
- Potentially hybrid search with BM25 + embeddings
Good options:
bge-large-en-v1.5e5-large-v2bge-m3if multilingual/heterogeneous corpus
If your corpus is multilingual
Use a multilingual model rather than translating everything:
bge-m3intfloat/multilingual-e5-large- OpenAI embeddings also handle multilingual reasonably well, but dedicated multilingual models are often stronger for retrieval in non-English corpora
2) Tradeoff: quality vs cost vs latency
Best quality
text-embedding-3-largebge-large-en-v1.5
Better cost/performance
text-embedding-3-smallbge-base-en-v1.5e5-base-v2
Self-hosting considerations
If you self-host embeddings:
- GPU/CPU inference cost
- Throughput
- Index rebuild time
- Dimensionality and vector store cost
3) Use one embedding model for retrieval, not multiple unless needed
Don’t overcomplicate at first. Pick one strong retriever embedding model and evaluate.
A common stack:
- Dense embeddings for semantic retrieval
- BM25 keyword retrieval
- Optional reranker on top
This is often better than trying to use multiple embedding models at once.
Chunking strategy: what works best
1) Chunk by structure first
Best practice is to split along:
- Markdown headings
- HTML sections
- PDF paragraphs
- Document sections/subsections
- Bullet lists
- Code blocks separately
Why this helps:
- Keeps related information together
- Makes retrieved context more coherent
- Reduces “orphaned” fragments
2) Recommended chunk sizes
A good starting range:
- Narrative/docs/help content: 300–600 tokens
- Technical docs / policies / manuals: 400–800 tokens
- FAQ-style content: often one Q/A per chunk, or slightly larger if answer spans multiple paragraphs
- Code/docs: chunk by function/class/module boundaries when possible
Overlap
- 10–20% overlap is usually enough
- Example: 500-token chunks with 50–100 token overlap
Too much overlap:
- Increases index size
- Creates near-duplicate chunks
- Can hurt diversity in retrieval results
Too little overlap:
- Risks splitting key context across chunks
3) Avoid purely fixed-size chunking unless necessary
Fixed-size chunking can split:
- A definition from its explanation
- A question from its answer
- A procedure step from its exception
- A code function from its docstring
Use fixed-size chunking only as a fallback.
4) Parent-child chunking is often excellent
This is a strong RAG pattern:
- Child chunks: small chunks embedded and retrieved
- Parent chunks: larger surrounding sections returned to the LLM
Benefits:
- High recall from small chunks
- Good context from larger chunks
Example:
- Embed 200–400 token chunks
- Store a mapping to 1,000–2,000 token parent sections
- Retrieve small chunks, then expand to parent section for generation
This works especially well for:
- Policies
- Technical documentation
- Legal content
- Long articles
Best chunking approach by content type
Documentation / knowledge base
Use:
- Heading-based splitting
- Paragraph-based splitting
- 400–800 token chunks
- 10–15% overlap
FAQs
Use:
- One question-answer per chunk
- Keep each answer intact if possible
- Add metadata with the question/topic
PDFs
PDFs are tricky because layout can be messy. Use:
- Layout-aware parsing
- Section/paragraph reconstruction
- Remove headers/footers
- Chunk by logical sections, not page breaks
Code repositories
Use:
- Function/class-level chunking
- Include surrounding docstrings/comments
- Possibly store file path, symbol name, and language metadata
Legal/policy docs
Use:
- Clause/section-based chunking
- Preserve numbering and headings
- Prefer smaller chunks with parent expansion
Retrieval quality tips beyond embeddings and chunking
1) Add metadata
Store metadata like:
- document title
- section heading
- URL/path
- timestamp/version
- language
- source type
This improves filtering and answer grounding.
2) Use hybrid retrieval
Dense embeddings are great, but keyword search helps with:
- exact names
- product IDs
- error codes
- legal terms
- acronyms
A strong baseline:
- BM25 + dense vectors
- Then rerank top results
3) Add a reranker if you can
A reranker can dramatically improve top-k quality.
Pipeline:
- Retrieve top 20–50 with embeddings/BM25
- Rerank with a cross-encoder or LLM-based reranker
- Pass top 3–8 chunks to the generation model
This often matters more than squeezing the last bit of performance from embeddings.
Practical starting configs
Option A: Strong general-purpose RAG
- Embeddings:
text-embedding-3-large - Chunking: 500 tokens, 75 token overlap
- Split on headings/paragraphs first
- Hybrid retrieval + reranker
Option B: Lower-cost production setup
- Embeddings:
text-embedding-3-small - Chunking: 400–600 tokens, 50–75 token overlap
- Hybrid retrieval
- Reranker if budget allows
Option C: Self-hosted open-source English
- Embeddings:
bge-large-en-v1.5ore5-large-v2 - Chunking: structure-aware, 400–700 tokens
- Hybrid retrieval + reranker
Option D: Multilingual
- Embeddings:
bge-m3 - Chunking: structure-aware, 400–700 tokens
- Hybrid retrieval
- Reranker with multilingual support
How to evaluate
Don’t choose blindly. Test on a small labeled set.
Measure:
- Recall@k: does the correct chunk appear in top k?
- MRR / nDCG: how high is the right chunk ranked?
- Answer accuracy: does final response actually answer correctly?
- Citation precision: are retrieved chunks relevant and well-grounded?
Create 20–100 real queries and inspect:
- retrieved chunks
- whether the answer is contained
- whether chunk boundaries are helping or hurting
A simple rule of thumb
If you’re unsure:
- Use structure-aware chunks
- Start around 500 tokens
- Overlap ~75 tokens
- Use
text-embedding-3-largeorbge-large-en-v1.5 - Add hybrid retrieval
- Add a reranker
That’s a very solid baseline.
If you want, I can also help you choose based on your exact setup:
- document type
- language(s)
- expected query types
- vector DB
- latency/cost constraints
If you share those, I can recommend a concrete embedding model + chunking configuration.