Long documents and big context
You want a model that holds a long document in its head, a contract, a codebase, a book, and reasons across all of it, instead of losing the thread partway through.
No board here measures this cleanly, so we are not naming a leader. Here is the yardstick and the dated anchor instead.
Long-context ability is read through LongBench v2, the one test built to require reasoning across a whole document rather than keyword matching.
The one test that requires reasoning across a whole long document, but its board is JavaScript-only with no readable current ranking.
No board here measures this cleanly, so we are not naming a winner. Judge it yourself on these things.
LongBench v2 is the right test, but there is no trustworthy current ranking to quote.
No clean answer yet: The official board is JavaScript-only and its readable scores are a stale early-2025 snapshot, so judge a model yourself on documents your own length.
- Does it keep facts straight across a long document, or contradict itself partway through?
- Does accuracy fall off a cliff once the document nears your real length, or stay steady?
A long context window is not the same as using it well: a model can advertise a huge window and still lose the thread in the middle, and a high score on a long-context test can mean the harness, not the model, did the reading. Test on your own document length before trusting a headline number.
The benchmark creates the need to know. The catalog explains the ideas behind it:
Human experts scored 53.7 percent under a 15-minute limit, the best directly-answering model 50.1 percent, and an extended-reasoning model reached 57.7 percent, just above the human number; blind guessing is about 25 percent (LongBench v2 paper, DEC 2024).
The only board is JavaScript-rendered and cannot be read from the page, and its readable figures are a stale early-2025 snapshot, so there is no trustworthy current ranking to quote.