Sign in
Back to tasks
Use case

Long documents and big context

You want a model that holds a long document in its head, a contract, a codebase, a book, and reasons across all of it, instead of losing the thread partway through.

Bottom line

No board here measures this cleanly, so we are not naming a leader. Here is the yardstick and the dated anchor instead.

The tests that matter

Long-context ability is read through LongBench v2, the one test built to require reasoning across a whole document rather than keyword matching.

LongBench v2

The one test that requires reasoning across a whole long document, but its board is JavaScript-only with no readable current ranking.

How to choose

No board here measures this cleanly, so we are not naming a winner. Judge it yourself on these things.

Reasoning across an entire long document

LongBench v2 is the right test, but there is no trustworthy current ranking to quote.

No clean answer yet: The official board is JavaScript-only and its readable scores are a stale early-2025 snapshot, so judge a model yourself on documents your own length.

  1. Does it keep facts straight across a long document, or contradict itself partway through?
  2. Does accuracy fall off a cliff once the document nears your real length, or stay steady?
What to watch

A long context window is not the same as using it well: a model can advertise a huge window and still lose the thread in the middle, and a high score on a long-context test can mean the harness, not the model, did the reading. Test on your own document length before trusting a headline number.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

LongBench v2
As of MAR 2025
What is known

Human experts scored 53.7 percent under a 15-minute limit, the best directly-answering model 50.1 percent, and an extended-reasoning model reached 57.7 percent, just above the human number; blind guessing is about 25 percent (LongBench v2 paper, DEC 2024).

Why there is no leader

The only board is JavaScript-rendered and cannot be read from the page, and its readable figures are a stale early-2025 snapshot, so there is no trustworthy current ranking to quote.

Backing board
No live board for this task.
Browse all 22 benchmarks