There is no trustworthy current leaderboard, so the most recent reliable figure is already months old.
LongBench v2
LongBench v2 tests whether a model can truly understand and reason over very long inputs, not just find one fact buried in them, using 503 hard multiple-choice questions over contexts from a few thousand up to two million words.
The trick is that answers require connecting information across the whole document, so a model cannot win by keyword-matching.
A model is given a long source, a multi-document set, a long dialogue history, or an entire code repository, and a four-option question whose answer depends on combining details from across it, under conditions where even human experts working under time pressure get only about half right.
Huge context-window numbers are a top marketing claim, and LongBench v2 tests whether a model actually uses that context to reason, which is what matters if you feed it long readings, contracts, or codebases. It is one of the benchmarks that is genuinely not solved yet, so it still separates models.
Accuracy on the four-option multiple-choice set (higher is better). Built by a Tsinghua-led team, released December 2024 and published at ACL 2025. The public board exists but is JavaScript-rendered and has not been kept current.
Blind guessing scores about 25 percent on the four-option format. Human experts reached only 53.7 percent under a 15-minute limit; the best plain model managed 50.1 percent, and o1-preview, using extended reasoning, reached 57.7 percent, just above the human number. That narrow spread is why this benchmark is considered unsaturated and hard.
There is no trustworthy current public top-3 for LongBench v2: the official board is JavaScript-rendered and its readable figures are a stale early-2025 snapshot, so this panel shows a dated anchor rather than a live ranking. Long-context scores are also sensitive to exactly how the context is fed in, so cross-model comparisons can be apples to oranges. Treat any single current number with caution.
The benchmark creates the need to know. The catalog explains the ideas behind it:
No reliable current leaderboard. The most cited figure is a dated anchor:
No reliable current public leaderboard: the official board is JavaScript-rendered and its readable scores are a stale early-2025 snapshot. The most-cited dated anchors are from the original paper (DEC 2024): human experts 53.7 percent under a 15-minute limit, the best directly-answering model 50.1 percent, and o1-preview 57.7 percent using extended reasoning; blind guessing is about 25 percent.