Top models score so close together that the gaps are inside the measurement noise, so this test no longer tells you which model is better.
AIME
AIME is one of the hardest math contests given to American high schoolers, repurposed as an AI test: 15 problems per exam, each answer a whole number from 0 to 999, with no partial credit and no multiple choice to guess from.
Top models now solve almost all of them.
A problem reads like a tough contest question, for example 'find the number of ordered integer pairs that satisfy the following system', with the answer a single integer from 0 to 999. There are no answer choices to guess among, and the problems are meant for pencil and paper, not a calculator.
AIME is a clean, fast-refreshing gauge of AI's step-by-step mathematical reasoning, the same machinery behind progress in science and engineering, and top models now beat the typical human qualifier. The reason to read it carefully is contamination: old exams leak into training data, so a sky-high score on a 2024 paper can be memorization, while performance on a freshly released exam is the honest signal.
Exact-match accuracy, the percent of problems solved (each model usually run several times and averaged). The trustworthy numbers come from MathArena, an independent academic group at ETH Zurich that runs every model itself on the newest exam, rather than from the contest organizers (the MAA), who only run the human competition.
Top human competitors solve about 7 to 10 of 15 problems (roughly 46 to 67 percent); random guessing is near zero because answers are integers. The 2025 wave of reasoning models pushed scores past 90 percent, and on the 2026 exam the leaders cluster between 93 and 97 percent.
Two cautions. Contamination is the big one: models reliably score 10 to 20 points higher on older AIME exams that have circulated online for years than on a freshly released one, so any 'AIME 2024' number is an unreliable upper bound (this page uses the fresh 2026 edition). Saturation is the other: the leaders are separated by a fraction of a point, well inside the margin of error, and the third-place slot is a three-way tie, so the ranking at the top is not a real ordering. Watch for scores quietly run with a calculator or code tool, which inflates a test meant for pencil and paper.
The benchmark creates the need to know. The catalog explains the ideas behind it:
The top ranks are a statistical tie. Read this as a cluster, not a clean number one.