Sign in
Benchmarks
Reasoning
Contested

The headline numbers are disputed, self-reported, or under revision, so treat the ranking as a claim, not a settled fact.

ARC-AGI v2

ARC-AGI v2 is the harder successor to ARC-AGI, a fresh set of novel grid puzzles rebuilt specifically to defeat the brute-force and memorization tricks that let models conquer the first version.

What it measures

Humans still solve these easily, and at its 2025 launch the best reasoning models scored only about 3 percent.

Example task

Like the original, each task shows a handful of before-and-after colored grids that share a hidden rule, then asks for the output on a new grid, but the rules are deliberately more compositional and novel so pattern-matching from training does not help. An average person solves roughly two thirds of them.

Why you should care

ARC-AGI-2 is the current front line of the argument over whether AI can really reason or just pattern-match, because it is built to be immune to memorization. It is also the sharpest example on this page of why provenance matters: the impressive public scores (one model is listed at 85 percent) are all self-reported lab claims that the benchmark's own foundation has not verified, and under the resource-constrained competition rules the best result is far lower.

How scoring works

Percent of 120 novel tasks solved, maintained by the ARC Prize Foundation, which separately runs a strict, resource-limited competition (no internet, fixed hardware, a time limit) for its grand prize. The public-leaderboard figures here come from third-party aggregators of lab-reported claims, not from ARC Prize's own verified testing.

How to read the numbers

Random chance is near zero; an average individual solves about 60 to 66 percent. At the May 2025 launch, leading reasoning models scored around 3 percent. The grand-prize bar is 85 percent on a hidden private set, which no team had reached as of mid-2026, and the best score under the constrained competition rules was only about 24 percent.

What to watch

Every number in this panel is a self-reported lab claim, not an ARC Prize verified result: the foundation's own tracker shows zero independently verified scores for the public ARC-AGI-2 leaderboard, and aggregators disagree on the ranking. The gap between the headline claims (a model listed at 85 percent) and the resource-constrained competition record (about 24 percent in 2025) is the whole story: raw, unconstrained scores and efficient, reproducible problem-solving are very different things. The verified high-water mark is lower and older (Gemini 3 Deep Think at 84.6 percent on a high-compute semi-private run that ARC Prize confirmed in February 2026), and the 85 percent grand prize on a hidden private set remains unclaimed.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders as of JUN 2026
GPT-5.5OpenAI
Self-reported
85.0%self-reported; not ARC Prize verified
APR 2026·benchlm.ai
GPT-5.4 ProOpenAI
Self-reported
83.3%self-reported lab claim
MAR 2026·benchlm.ai
Gemini 3.1 ProGoogle DeepMind
Self-reported
77.1%self-reported lab claim
JUN 2026·benchlm.ai
StatusContested
Board last updated JUN 2026, may be behind.
View the leaderboard