AI & Data › LLM & AI Engineering
Reranking
Re-ordering retrieved results with a stronger model.
Also known as: reranking, re-ranking, result reranking
Reranking is a second pass over the candidates a retriever returned. A first stage is fast but approximate, pulling in a few hundred possibly relevant passages; a reranker then scores each candidate against the query more carefully and reorders them, so the best few reach the model or the user.
query → fast retrieval (many candidates) → reranker scores each (query, candidate) → top few
The two stages trade speed for precision. Retrieval is cheap enough to run over a large index; reranking is more expensive per item but only sees a short list. Separating them keeps latency acceptable while improving what ends up in context.
The classic mistakes:
- Reranking too many candidates. Scoring thousands of items per query erases the latency gain. Keep the candidate list short and bounded.
- Skipping evaluation. A reranker that looks sophisticated can still hurt results on your queries. Compare with and without it.
- Assuming it fixes bad retrieval. If the relevant passage is absent from candidates, reranking cannot recover it. Measure recall at the first stage.
- Ignoring the cost budget. Reranking adds per-query compute or calls. Account for it in latency and price.
- Using it everywhere. Simple lookups with exact identifiers often need no reranker.
Practice: retrieve broadly, rerank a short list, and confirm the gain on a labelled query set.
Watch latency at the tail, not just the average. A reranker that adds a modest median cost can still produce slow outliers on large candidate lists, and those outliers are what users complain about first.