Sign in
Benchmarks
Long context
Dated

There is no trustworthy current leaderboard, so the most recent reliable figure is already months old.

LongBench v2

LongBench v2 tests whether a model can truly understand and reason over very long inputs, not just find one fact buried in them, using 503 hard multiple-choice questions over contexts from a few thousand up to two million words.

What it measures

The trick is that answers require connecting information across the whole document, so a model cannot win by keyword-matching.

Example task

A model is given a long source, a multi-document set, a long dialogue history, or an entire code repository, and a four-option question whose answer depends on combining details from across it, under conditions where even human experts working under time pressure get only about half right.

Why you should care

Huge context-window numbers are a top marketing claim, and LongBench v2 tests whether a model actually uses that context to reason, which is what matters if you feed it long readings, contracts, or codebases. It is one of the benchmarks that is genuinely not solved yet, so it still separates models.

How scoring works

Accuracy on the four-option multiple-choice set (higher is better). Built by a Tsinghua-led team, released December 2024 and published at ACL 2025. The public board exists but is JavaScript-rendered and has not been kept current.

How to read the numbers

Blind guessing scores about 25 percent on the four-option format. Human experts reached only 53.7 percent under a 15-minute limit; the best plain model managed 50.1 percent, and o1-preview, using extended reasoning, reached 57.7 percent, just above the human number. That narrow spread is why this benchmark is considered unsaturated and hard.

What to watch

There is no trustworthy current public top-3 for LongBench v2: the official board is JavaScript-rendered and its readable figures are a stale early-2025 snapshot, so this panel shows a dated anchor rather than a live ranking. Long-context scores are also sensitive to exactly how the context is fed in, so cross-model comparisons can be apples to oranges. Treat any single current number with caution.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders

No reliable current leaderboard. The most cited figure is a dated anchor:

No reliable current public leaderboard: the official board is JavaScript-rendered and its readable scores are a stale early-2025 snapshot. The most-cited dated anchors are from the original paper (DEC 2024): human experts 53.7 percent under a 15-minute limit, the best directly-answering model 50.1 percent, and o1-preview 57.7 percent using extended reasoning; blind guessing is about 25 percent.

StatusDated
View the board