Skip to main content
Updated July 12, 2026
A reranker reorders the chunks your retriever returns so the best ones land on top - the cheapest upgrade to RAG accuracy. But the highest-scoring models often ship noncommercial weights, so “open” rarely means self-hostable. We ranked 12 on quality, price, speed, and license.

Best Reranker Models


Zerank 2

Top-accuracy multilingual RAG
The most accurate reranker in the current benchmark, and among the fastest and cheapest too - if you can live with weights you can’t ship commercially.
Score 100%Price License Open weightLatency 265 ms
  • It leads on ranking quality while staying near the front on speed, a rare combination. Its relevance scores are well calibrated, so you can set real cutoff thresholds instead of guessing.
  • Instruction-following and genuine 100+ language coverage make it the strongest pick for multilingual or domain-specialized retrieval.
  • The weights are noncommercial, so self-hosting in a product needs a paid ZeroEntropy license, and most teams land on the metered API anyway.
  • Local runs also need a high-end GPU. For weights you can actually ship, Zerank 1 Small or Qwen3 are the alternatives.
Quality-first enterprise RAG
Cohere’s v4 flagship is the strongest proprietary reranker here, a real jump over 3.5 that shines on long, entity-heavy enterprise documents.
Score 97%Price License ProprietaryLatency 614 ms
  • It covers 100+ languages and jumped to roughly 32K context, so long filings and reports rerank without the pre-chunking dance.
  • Quality is consistent across domains with the biggest gains on finance, business, and entity-heavy content, and it sits near the very top on preference-based evaluation.
  • Closed weights mean no self-hosting, though Cohere does offer private managed deployment if data residency is the concern.
  • Latency is middle-of-the-pack, slower than Zerank 2 and its own Fast tier, and there’s no instruction-based steering - if you want that, Voyage 2.5 is the closer fit.
Balanced instruction-following RAG
Voyage’s generalist reranker is the balanced pick, and the only strong proprietary option here you can steer with plain-language instructions.
Score 70%Price License ProprietaryLatency 613 ms
  • Natural-language instructions let you steer ranking - emphasize a field, prefer a document type, disambiguate a query - a real edge for agentic and conversational retrieval.
  • It also leads the proprietary set on pure retrieval-accuracy metrics and handles 32K context, so it’s a safe balanced default.
  • It’s API-only from a single vendor, with no local or private-deployment route, unlike Cohere. Language coverage is narrower than Cohere’s, and on preference-based ranking it sits below Cohere Pro and Zerank 2.
  • It wins on balance, not on any single number.
Self-hostable lightweight reranker
The previous-generation ZeroEntropy small model earns its spot on this list for one reason: permissive weights you can actually deploy.
Score 68%Price License Open weightLatency 248 ms
  • Apache 2.0 weights and a small footprint mean it drops into a commercial product with no licensing conversation and runs on ordinary hardware.
  • It’s the fastest model in the Zerank line, cheap to self-host, and punches above its size on quality - exactly what Zerank 2’s license won’t let you do.
  • It’s clearly below Zerank 2, Cohere 4, and Voyage 2.5 on ranking quality - this is the “good enough and yours” option, not the accuracy leader.
  • It’s English-centric with no instruction-following, so for multilingual or steerable ranking you want Zerank 2 or an open Qwen3.
Cost-efficient high-volume reranking
The cheaper Voyage tier keeps the instruction-following and long context of 2.5, giving up a little quality for a much lower price.
Score 62%Price License ProprietaryLatency 616 ms
  • It carries the full 2.5 feature set - instruction-following, 32K context, multilingual - into a much cheaper tier, which makes it the value pick when query volume is high.
  • Quality holds up better than the price suggests, landing above Cohere’s v4 variants on pure retrieval accuracy.
  • There’s a real if small quality step-down from full 2.5, so skip it when ranking quality is the priority. And despite the “Lite” name it isn’t faster; latency matches 2.5, so the only reason to choose it over 2.5 is cost.
  • Same API-only, single-vendor limits.
Low-latency enterprise reranking
The speed-tuned v4 tier is faster than Pro and keeps the same context and languages, but it’s a specialized tool, not a universal upgrade.
Score 59%Price License ProprietaryLatency 447 ms
  • It’s meaningfully faster than Pro with higher throughput, the one to reach for when your latency budget is tight.
  • You keep the same 100+ language coverage and 32K context, and on enterprise content - finance, business, entity-heavy queries - it still improves on the older 3.5.
  • The catch is uneven quality: on argumentation-heavy and general web-style questions it can fall behind the older 3.5. It ranks below both Voyage 2.5 tiers, and like all Cohere models it’s closed-weight.
  • Pick it for speed on enterprise content, not as a blanket upgrade.
Top-quality open-weight reranking
The largest open Qwen3 reranker offers frontier-adjacent quality under a truly permissive license, but its latency makes it a batch tool, not a live one.
Score 47%Price License Open weightLatency 4,687 ms
  • Apache 2.0 gives you unrestricted commercial use at a quality tier where that’s rare, plus full on-prem control. It covers 100+ languages including code, takes task instructions, handles 32K context, and tops academic multilingual retrieval benchmarks.
  • If sovereignty and commercial freedom both matter, it has few peers.
  • The dealbreaker is speed - the slowest model here, pushing it to offline or batch reranking.
  • Self-hosting needs a high-end GPU, and its quality edge is benchmark-dependent: it tops academic multilingual tests but trails Zerank 2 and Cohere on preference ranking.
Instruction-steered enterprise RAG
This is the reranker to reach for when your corpus has conflicting sources and you need to steer ranking by recency, authority, or document type.
Score 46%Price License Open weightLatency 3,333 ms
  • It’s purpose-built for instruction steering: a plain-language instruction can prioritize recent documents, trusted internal sources, or a specific document type - useful when relevance alone can’t settle contradictions between sources.
  • Multilingual coverage spans 100+ languages, context runs to 32K, and it does this in a compact 2B model.
  • The weights are noncommercial with share-alike terms, so commercial use routes you to the paid API, and self-hosting is slow on high-end hardware.
  • Instruction steering only earns its keep if you need cross-source arbitration - for plain relevance reranking, Voyage 2.5 is faster and less restricted.
Permissive open-weight baseline
The reranker most RAG stacks ship by default - free, permissive, and multilingual - now an aging baseline that newer open models beat on quality.
Score 0%Price License Open weightLatency 2,383 ms
  • Apache 2.0 makes it free to self-host commercially with no asterisks, and it’s small enough to run on a typical machine, even CPU.
  • Multilingual coverage is proven across 100+ languages, and it’s integrated into nearly every RAG framework, so it’s the safe, known-quantity starting point.
  • It’s an older baseline, not a frontier model, and Qwen3’s open rerankers beat it on accuracy while staying just as permissive. Long documents are a weak spot: it was tuned for short passages and quietly truncates long chunks unless you raise the limit.
  • It’s also slow.
Fast long-context reranking
The newest Jina reranker is the fastest here and handles the longest documents, thanks to a new listwise design, if you can accept noncommercial weights.
Score Not rankedPrice License Open weightLatency 167 ms
  • Its listwise design reranks the whole candidate set in one pass instead of scoring documents one by one - the reason it’s the fastest model here and a clear step up from v2.
  • It also handles the longest context in this list and runs on a typical machine.
  • The weights are noncommercial, so shipping it in a commercial product means Jina’s paid API - the same catch as Zerank 2 and Contextual.
  • It also sits outside our leaderboard, so its quality case rests on Jina’s own benchmarks rather than head-to-head results.
Cheap fast local reranking
The smallest Qwen3 reranker is the cheapest, most deployable option here - permissive weights that run fast on ordinary hardware, with a lower quality ceiling.
Score Not rankedPrice License Open weightLatency 445 ms
  • It inherits the Apache 2.0 license, 100+ languages, and 32K context of the larger Qwen3 rerankers, but runs on a typical machine and reranks fast enough for live use.
  • It’s the cheapest hosted option and a sensible default when cost, latency, and commodity hardware matter more than peak accuracy.
  • The ceiling is real: on hard multi-hop questions or nuanced relevance it noticeably trails the 8B and proprietary leaders. This isn’t a quality play - it competes on price, speed, and license.
  • If accuracy is the bottleneck, step up to Qwen3 4B or a paid API.
Cross-lingual retrieval reranking
NVIDIA’s 1B reranker is fast and strong cross-lingually, but it’s built to run as a GPU microservice, which narrows who can realistically use it.
Score Not rankedPrice License Open weightLatency 223 ms
  • Its standout is cross-lingual retrieval - evaluated across 26 languages with strong results when query and document languages differ, plus solid long-document recall.
  • It’s genuinely fast, and unlike the noncommercial open models here its weights carry commercial-friendly terms, so you can actually ship it.
  • The supported path needs recent NVIDIA GPUs, so it’s a non-starter on CPU or other hardware - the raw weights run elsewhere but unoptimized.
  • There’s no public per-token price, so cost is infrastructure-based and hard to compare, and context tops out at 8K, the shortest here.

How to Choose

When choosing between these models, consider:
  • Access: Decide first whether you’ll call a hosted API, use a managed platform, or self-host. Most of the strongest models are API-first; only some open-weight options are realistic to run yourself, and a few of those need a high-end GPU. That one choice drives cost, privacy, latency, and setup work.
  • Quality: We use the Agentset Rerankers Leaderboard as the score, normalized to 0-100%. It ranks rerankers by head-to-head Elo from preference judgments on real retrieval tasks - a better proxy for “did it put the right chunk on top” than a single accuracy metric. A 0% is the bottom of the measured range, not a broken model, and three highlighted picks (Jina v3, Qwen3 0.6B, Nemotron) aren’t on the board yet.
  • Price: We normalize to USD per 1M reranked tokens for one clean axis. Watch the fine print: Cohere and Voyage bill per search or request natively, so their per-token figures are conversions, and NVIDIA’s Nemotron has no public token price at all.
  • Reranking Latency: Treat the millisecond figures as directional. Most come from Agentset’s hosted top-50 benchmark, but Jina v3, Qwen3 0.6B, and Nemotron use a different exact-model GPU benchmark, so they aren’t strictly comparable. Use latency mainly to separate “fast enough for live chat” from “batch only” - Qwen3 8B and Contextual v2 are firmly in the second group.

Other Models We Considered

Zerank 1 (ZeroEntropy) — Still-strong prior flagship, but Zerank 2 wins at the same price.Cohere Rerank 3.5 (Cohere) — A common production baseline; Rerank 4 adds much longer context.Jina Reranker v2 Base Multilingual (Jina AI) — Compact multilingual predecessor, now superseded by the faster v3.Qwen3 Reranker 4B (Qwen) — The middle size, a quality-speed compromise between 0.6B and 8B.mxbai-rerank-large-v2 (Mixedbread) — Permissive multilingual model with code retrieval, worth testing for coding RAG.GTE Reranker ModernBERT Base (Alibaba-NLP) — Tiny English reranker with long context and permissive weights.MS MARCO MiniLM L6 v2 (Sentence Transformers) — The classic tiny English baseline older RAG tutorials default to.

Frequently Asked Questions

Zerank 2. It tops our leaderboard on ranking quality while staying among the fastest and cheapest hosted options. The catch is licensing: its open weights are noncommercial, so most teams use its metered API rather than self-hosting. If you want a fully commercial, closed managed service instead, Cohere Rerank 4 Pro is the closest rival.
For most RAG pipelines, Voyage Rerank 2.5 or Cohere Rerank 4 Pro are the safe managed defaults - high quality, long context, and no infrastructure to run. If cost matters more than the last few points of accuracy, Voyage 2.5 Lite and Qwen3 Reranker 0.6B are strong value picks.
For permissive, ship-it-anywhere weights, Qwen3 Reranker (0.6B on typical hardware, 8B if you have a GPU and can accept high latency) and BGE Reranker v2 M3 are the cleanest choices, all Apache 2.0. Zerank 1 Small is the fast, small option under the same terms. Watch out: several “open” rerankers, including Zerank 2, Jina v3, and Contextual v2, are noncommercial.
Usually, yes - reranking is often the cheapest way to lift answer quality, because it fixes the order of what you already retrieved. But it only helps when the right chunk is somewhere in your top results and just ranked too low. If recall is bad and the right chunk isn’t retrieved at all, fix retrieval first; a reranker can’t surface what isn’t there.
It depends on the metric. Cohere Rerank 4 Pro leads on preference-based ranking and covers more languages; Voyage 2.5 leads on pure retrieval-accuracy metrics and adds instruction-following, which Cohere lacks. Pick Cohere for broad multilingual enterprise content, Voyage when you want to steer ranking with instructions. They’re close enough to test both on your data.
No, and this is the biggest trap in the category. Several top open-weight rerankers - Zerank 2, Jina Reranker v3, Contextual v2 - ship under noncommercial licenses, so using the weights in a product needs a paid agreement. For unrestricted commercial self-hosting, stick to Apache 2.0 models like Qwen3, BGE v2 M3, and Zerank 1 Small.
Roughly, but not perfectly. Leaderboard rank tells you which models are contenders, yet the order shifts with your domain, language, and document length - and a model that tops academic tests, like Qwen3 8B, can land mid-pack on preference-based ranking. Treat the score as a shortlist filter, then measure your top two or three on your own queries.