Prompt
need embeddings for RAG with chunking and re-ranking
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For RAG, a solid setup is:
- Chunk the documents
- Create embeddings for each chunk
- Store them in a vector database
- Retrieve top-k chunks for a query
- Re-rank the retrieved chunks with a cross-encoder or reranker
- Pass the best chunks to the LLM
Recommended approach
1) Chunking
Use chunks that are:
- Small enough to capture a single idea
- Large enough to preserve context
Typical starting point:
- Chunk size: 300–800 tokens
- Overlap: 10–20%
Best practice:
- Split by structure first: headings, paragraphs, sections
- Then apply token-based chunking
- Keep metadata:
- document_id
- section title
- chunk index
- source URL / page number
2) Embeddings
Use a strong embedding model. Good choices depend on your stack:
- OpenAI:
text-embedding-3-smallortext-embedding-3-large - Open-source:
bge-large-ene5-largegte-large- multilingual variants if needed
If your corpus is domain-specific, consider:
- fine-tuning embeddings
- or at least testing several models on a retrieval benchmark
3) Vector store
Store chunk embeddings in:
- FAISS
- Pinecone
- Weaviate
- Milvus
- pgvector
- Qdrant
4) Retrieval
At query time:
- embed the query
- retrieve top 20–50 chunks by cosine similarity / dot product
5) Re-ranking
Use a reranker to improve relevance.
Good rerankers:
bge-reranker-large- Cohere rerank models
- cross-encoders from sentence-transformers
Pipeline:
- initial vector search: top 20–50
- rerank these with a cross-encoder
- keep top 3–10 for generation
This usually improves precision a lot.
Simple recommended pipeline
If you want a practical default:
- Chunk size: 500 tokens
- Overlap: 100 tokens
- Embedding model:
text-embedding-3-smallorbge-large-en - Retrieve top 30
- Re-rank to top 5
- Feed top 5 chunks to the LLM
Example architecture
Documents
-> parse + clean
-> chunk
-> embed each chunk
-> store in vector DB
User query
-> embed query
-> retrieve top-k chunks
-> rerank with cross-encoder
-> send best chunks to LLM
Important tips
- Don’t chunk too small, or you lose context
- Don’t chunk too large, or retrieval becomes noisy
- Re-ranking is especially useful when:
- queries are vague
- documents are long
- you have many similar chunks
- If the documents are code, legal, or scientific text, tune chunking differently
Example chunking strategy
For markdown or docs:
- split by headings
- keep section text together
- if a section is too long, split by paragraph
- if still too long, split by token window
If you want the best practical setup
Use:
- Hierarchical chunking
- Dense embeddings
- Hybrid retrieval: vector + keyword/BM25
- Re-ranking
- Context compression before the LLM if needed
That combination is usually stronger than embeddings alone.
If you want, I can also give you:
- a Python code example for chunking + embeddings + reranking, or
- a recommended model shortlist for your use case.