Sign in
Benchmarks
Reasoning
Near ceiling(near tie)

Top models score so close together that the gaps are inside the measurement noise, so this test no longer tells you which model is better.

ARC-AGI v1

ARC-AGI v1 is a set of colored-grid puzzles where you see a few examples of a transformation and must infer the rule, then apply it to a new grid.

What it measures

It tests whether a model can learn a brand-new pattern on the spot rather than recall a trained one. It was meant to be easy for humans and hard for AI, and for years it was.

Example task

You are shown three or four small before-and-after grids of colored squares that share a hidden rule (say, the shape gets reflected and recolored), then a new input grid, and you must draw the correct output. People find these easy; until recently, AI could barely do them.

Why you should care

ARC-AGI was the famous test that machines could not crack: for years top models scored near zero while ordinary people solved the puzzles easily, making it a favorite argument that AI lacked real reasoning. The reason to care now is the reversal: by 2026 the best systems score 96 to 98 percent, matching the human reference, so this particular challenge has essentially been beaten, and its successor ARC-AGI-2 has taken its place.

How scoring works

Accuracy on a held-out (semi-private) set, run and verified by the ARC Prize Foundation, which also tracks the compute cost per task to discourage winning by brute force. The original December 2024 breakthrough (OpenAI's o3 at 75.7 percent) is the famous milestone; the board has since climbed to the ceiling.

How to read the numbers

Random guessing is near zero. The human references on the same board are about 77 percent for an average crowd worker and 98 percent for a panel. The 2024 breakthrough was o3 at 75.7 percent within the cost limit, and by mid-2026 the leaders sit at 96 to 98 percent, so v1 is effectively saturated.

What to watch

ARC-AGI v1 is essentially solved, so the ordering at the very top is a near-tie among models clustered between 96 and 98 percent, right at the human-panel reference. Two cautions remain. The scores vary widely in cost (the top three here range from about 50 cents to over 7 dollars per task), so a higher score can simply mean more compute was spent, which is why cost is shown next to each number. And critics have always disputed the AGI framing even while accepting the numbers. The real action has moved to ARC-AGI-2.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders as of JUN 2026
Gemini 3.1 Pro (Preview)Google
98% (ARC-AGI-1 Semi-Private eval, $0.52/task)human panel reference is 98%
JUN 2026·arcprize.org
GPT-5.5 Pro (High)OpenAI
96.5% (ARC-AGI-1 Semi-Private eval, $4.53/task)note the higher cost per task
JUN 2026·arcprize.org
Gemini 3 Deep Think (2/26)Google
96% (ARC-AGI-1 Semi-Private eval, $7.17/task)ceiling cluster, not a clear gap
JUN 2026·arcprize.org

The top ranks are a statistical tie. Read this as a cluster, not a clean number one.

StatusNear ceiling
Board last updated JUN 2026, may be behind.
View the leaderboard