Sign in
Benchmarks
Math
Near ceiling(near tie)

Top models score so close together that the gaps are inside the measurement noise, so this test no longer tells you which model is better.

AIME

AIME is one of the hardest math contests given to American high schoolers, repurposed as an AI test: 15 problems per exam, each answer a whole number from 0 to 999, with no partial credit and no multiple choice to guess from.

What it measures

Top models now solve almost all of them.

Example task

A problem reads like a tough contest question, for example 'find the number of ordered integer pairs that satisfy the following system', with the answer a single integer from 0 to 999. There are no answer choices to guess among, and the problems are meant for pencil and paper, not a calculator.

Why you should care

AIME is a clean, fast-refreshing gauge of AI's step-by-step mathematical reasoning, the same machinery behind progress in science and engineering, and top models now beat the typical human qualifier. The reason to read it carefully is contamination: old exams leak into training data, so a sky-high score on a 2024 paper can be memorization, while performance on a freshly released exam is the honest signal.

How scoring works

Exact-match accuracy, the percent of problems solved (each model usually run several times and averaged). The trustworthy numbers come from MathArena, an independent academic group at ETH Zurich that runs every model itself on the newest exam, rather than from the contest organizers (the MAA), who only run the human competition.

How to read the numbers

Top human competitors solve about 7 to 10 of 15 problems (roughly 46 to 67 percent); random guessing is near zero because answers are integers. The 2025 wave of reasoning models pushed scores past 90 percent, and on the 2026 exam the leaders cluster between 93 and 97 percent.

What to watch

Two cautions. Contamination is the big one: models reliably score 10 to 20 points higher on older AIME exams that have circulated online for years than on a freshly released one, so any 'AIME 2024' number is an unreliable upper bound (this page uses the fresh 2026 edition). Saturation is the other: the leaders are separated by a fraction of a point, well inside the margin of error, and the third-place slot is a three-way tie, so the ranking at the top is not a real ordering. Watch for scores quietly run with a calculator or code tool, which inflates a test meant for pencil and paper.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders as of MAY 2026
Step-3.5-FlashStepFun AI
96.67% +/- 3.21%MathArena, fresh 2026 exam
MAY 2026·matharena.ai
Kimi-K2.6Moonshot AI
96.4%0.27 points behind #1
MAY 2026·huggingface.co
Kimi-K2.5Moonshot AI
95.83% +/- 3.58%three-way tie for third
MAY 2026·matharena.ai

The top ranks are a statistical tie. Read this as a cluster, not a clean number one.

StatusNear ceiling
Board last updated JUN 2026, may be behind.
View the leaderboard