Prompt
How do I re-rank retrieved passages before sending them to the model?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
A good way to re-rank retrieved passages is to add a second-stage ranker after your initial retriever, before you build the prompt for the model.
Common pipeline
- Retrieve a larger pool of candidates
- e.g. top 20–100 passages from BM25, vector search, or hybrid retrieval.
- Re-rank those candidates with a more precise scoring method.
- Select top-k passages and send only those to the model.
Re-ranking methods
1) Cross-encoder re-ranker
This is the most common and usually the best quality option.
- Input:
(query, passage)pair - Output: relevance score
- The model reads the query and passage together, so it’s more accurate than embeddings alone.
Example flow
candidates = retriever.search(query, top_k=50)
scored = [(passage, reranker.score(query, passage)) for passage in candidates]
ranked = sorted(scored, key=lambda x: x[1], reverse=True)
top_passages = [p for p, s in ranked[:5]]
Pros
- High relevance accuracy
- Great for final selection
Cons
- Slower than retrieval
- More expensive at scale
2) LLM-based re-ranking
Use an LLM to compare or score passages for the query.
Best for
- Small candidate sets
- High-value queries
- Situations where you want reasoning-based ranking
Example prompt
Rank these passages by how well they answer the query. Return the top 5 with scores and a short reason.
Pros
- Can understand nuanced relevance
- Flexible
Cons
- More expensive
- Less deterministic
- Harder to scale
3) Lightweight heuristic re-ranking
Useful when you want a fast boost without another model.
Possible signals:
- BM25 score
- Vector similarity score
- Keyword overlap
- Passage length penalty
- Source freshness or authority
- Metadata filters
You can combine them:
final_score = (
0.5 * vector_score +
0.3 * bm25_score +
0.2 * metadata_boost
)
Practical recommendation
For most RAG systems:
- Stage 1: retrieve top 50–100 with vector or hybrid search
- Stage 2: re-rank with a cross-encoder
- Stage 3: send top 3–10 passages to the model
This usually gives the best balance of quality and latency.
Tips
- Re-rank before deduplication if passages may be near-duplicates, then remove duplicates after scoring.
- Consider chunk-level and document-level ranking separately.
- Keep the total context budget in mind; don’t just send the highest scores if they are redundant.
- If passages are long, score multiple chunks and then aggregate by document.
If you want a concrete implementation
I can show you:
- a Python example with SentenceTransformers cross-encoder
- an OpenAI/LLM-based re-ranker prompt
- or a full RAG pipeline with hybrid retrieval + re-ranking