The headline numbers are disputed, self-reported, or under revision, so treat the ranking as a claim, not a settled fact.
ARC-AGI v2
ARC-AGI v2 is the harder successor to ARC-AGI, a fresh set of novel grid puzzles rebuilt specifically to defeat the brute-force and memorization tricks that let models conquer the first version.
Humans still solve these easily, and at its 2025 launch the best reasoning models scored only about 3 percent.
Like the original, each task shows a handful of before-and-after colored grids that share a hidden rule, then asks for the output on a new grid, but the rules are deliberately more compositional and novel so pattern-matching from training does not help. An average person solves roughly two thirds of them.
ARC-AGI-2 is the current front line of the argument over whether AI can really reason or just pattern-match, because it is built to be immune to memorization. It is also the sharpest example on this page of why provenance matters: the impressive public scores (one model is listed at 85 percent) are all self-reported lab claims that the benchmark's own foundation has not verified, and under the resource-constrained competition rules the best result is far lower.
Percent of 120 novel tasks solved, maintained by the ARC Prize Foundation, which separately runs a strict, resource-limited competition (no internet, fixed hardware, a time limit) for its grand prize. The public-leaderboard figures here come from third-party aggregators of lab-reported claims, not from ARC Prize's own verified testing.
Random chance is near zero; an average individual solves about 60 to 66 percent. At the May 2025 launch, leading reasoning models scored around 3 percent. The grand-prize bar is 85 percent on a hidden private set, which no team had reached as of mid-2026, and the best score under the constrained competition rules was only about 24 percent.
Every number in this panel is a self-reported lab claim, not an ARC Prize verified result: the foundation's own tracker shows zero independently verified scores for the public ARC-AGI-2 leaderboard, and aggregators disagree on the ranking. The gap between the headline claims (a model listed at 85 percent) and the resource-constrained competition record (about 24 percent in 2025) is the whole story: raw, unconstrained scores and efficient, reproducible problem-solving are very different things. The verified high-water mark is lower and older (Gemini 3 Deep Think at 84.6 percent on a high-compute semi-private run that ARC Prize confirmed in February 2026), and the 85 percent grand prize on a hidden private set remains unclaimed.
The benchmark creates the need to know. The catalog explains the ideas behind it: