Sign in
Back to tasks
Use case

Reasoning on genuinely new problems

You want a model that learns a brand-new rule on the spot and applies it, the skill that separates real reasoning from sophisticated pattern-matching.

Bottom line

On ARC-AGI v2, the test rebuilt to defeat memorization, every headline score in the public ranking is a self-reported lab claim the benchmark's own foundation has not verified. As of JUN 2026 the honest read is that there is no verified leader: the foundation's tracker shows zero independently confirmed public scores.

The tests that matter

Novel reasoning is read through ARC-AGI v2, the current and unsaturated front line, with ARC-AGI v1 as the now-solved predecessor.

ARC-AGI v2

Built to be immune to memorization, but every public score is a self-reported lab claim, not an ARC Prize verified result.

ARC-AGI v1

The famous test machines could not crack, now effectively solved at 96 to 98 percent, so it no longer separates the top.

How to choose
Solving genuinely new, compositional puzzles

ARC-AGI v2 is the right test, but treat the public leaderboard as unverified claims; the only trustworthy figures are the ARC Prize foundation's own confirmed runs and the constrained competition record.

ARC-AGI v2
The classic on-the-spot pattern test

ARC-AGI v1 is essentially beaten, so a top score here is table stakes, not a differentiator. The action has moved to v2.

ARC-AGI v1
What to watch

This is the sharpest example on the surface of why provenance matters: the impressive public numbers are self-reported and unverified, and the gap between a headline claim near 85 percent and the verified, resource-constrained record near 24 percent is the whole story. Raw, unconstrained scores and efficient, reproducible problem-solving are very different things.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

ARC-AGI v2
As of JUN 2026
Best on novel puzzles built to defeat memorization
GPT-5.5OpenAI
Self-reported
85.0%self-reported; not ARC Prize verified
APR 2026ยทbenchlm.ai
Caveat

Every number in the public ARC-AGI v2 ranking is a self-reported lab claim, not an ARC Prize verified result, and aggregators disagree on the order. Under the strict, resource-limited competition the best result was only about 24 percent, far below the headline claims.