Reasoning on genuinely new problems
You want a model that learns a brand-new rule on the spot and applies it, the skill that separates real reasoning from sophisticated pattern-matching.
On ARC-AGI v2, the test rebuilt to defeat memorization, every headline score in the public ranking is a self-reported lab claim the benchmark's own foundation has not verified. As of JUN 2026 the honest read is that there is no verified leader: the foundation's tracker shows zero independently confirmed public scores.
Novel reasoning is read through ARC-AGI v2, the current and unsaturated front line, with ARC-AGI v1 as the now-solved predecessor.
Built to be immune to memorization, but every public score is a self-reported lab claim, not an ARC Prize verified result.
The famous test machines could not crack, now effectively solved at 96 to 98 percent, so it no longer separates the top.
ARC-AGI v2 is the right test, but treat the public leaderboard as unverified claims; the only trustworthy figures are the ARC Prize foundation's own confirmed runs and the constrained competition record.
ARC-AGI v2ARC-AGI v1 is essentially beaten, so a top score here is table stakes, not a differentiator. The action has moved to v2.
ARC-AGI v1This is the sharpest example on the surface of why provenance matters: the impressive public numbers are self-reported and unverified, and the gap between a headline claim near 85 percent and the verified, resource-constrained record near 24 percent is the whole story. Raw, unconstrained scores and efficient, reproducible problem-solving are very different things.
The benchmark creates the need to know. The catalog explains the ideas behind it:
Every number in the public ARC-AGI v2 ranking is a self-reported lab claim, not an ARC Prize verified result, and aggregators disagree on the order. Under the strict, resource-limited competition the best result was only about 24 percent, far below the headline claims.