Sign in
Back to tasks
Use case

Math and step-by-step reasoning

You want a model that reasons through a multi-step problem to the right final number, on problems it could not have seen before, not one that pattern-matches a memorized solution.

Bottom line

On the freshest contest exam, AIME 2026, the leaders cluster between about 94 and 97 percent, separated by a fraction of a point, so the very top is inside the noise. As of JUN 2026 the cleanest read is that several models are effectively tied; do not over-read whoever is listed first.

The tests that matter

Math is read through AIME for contest-style reasoning on a fresh exam, and FrontierMath for research-level problems, with the latter currently under an error review.

AIME

Independent and refreshed on each new exam, but the leaders are nearly tied and old editions leak into training.

FrontierMath

Research-level and guess-proof, but scores are pre-correction while the maintainer reviews about a third of the problems.

How to choose
Contest-style problems with exact integer answers

AIME on the fresh 2026 exam is the honest signal. Ignore any AIME 2024 number, which is inflated by contamination, and watch for scores quietly run with a calculator or code tool.

AIME
Genuinely novel, research-level mathematics

FrontierMath is the only test at this difficulty, but read it as provisional while the error review and the funding conflict are unresolved.

FrontierMath
What to watch

Math is where contamination bites hardest. A model can look brilliant on an old exam it memorized and ordinary on a fresh one, so always ask which exam a number came from. At the very top, the leaders are usually a fraction of a point apart, which is a tie, not a ranking.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

AIME
As of MAY 2026
Best for contest-style math on fresh problems
Step-3.5-FlashStepFun AI
96.67% +/- 3.21%MathArena, fresh 2026 exam
MAY 2026·matharena.ai
Caveat

AIME leaders are separated by a fraction of a point, well inside the margin of error, so read the top as a cluster, not a ranking. Scores on older AIME papers run 10 to 20 points higher because those exams leaked into training; this pick uses the fresh 2026 exam.

Best for research-level mathematics
Claude Fable 5 (max)Anthropic
87.8% +/- 5.2 (FrontierMath Tier 4 v2)Epoch eval; experts about 19% on easier tiers
JUN 2026·lmcouncil.ai
Caveat

FrontierMath numbers are contested: Epoch AI found fatal errors in roughly one third of the problems, so every current score is pre-correction, and OpenAI funded the benchmark with access to most problems. Read the standings as provisional.