Math and step-by-step reasoning
You want a model that reasons through a multi-step problem to the right final number, on problems it could not have seen before, not one that pattern-matches a memorized solution.
On the freshest contest exam, AIME 2026, the leaders cluster between about 94 and 97 percent, separated by a fraction of a point, so the very top is inside the noise. As of JUN 2026 the cleanest read is that several models are effectively tied; do not over-read whoever is listed first.
Math is read through AIME for contest-style reasoning on a fresh exam, and FrontierMath for research-level problems, with the latter currently under an error review.
Independent and refreshed on each new exam, but the leaders are nearly tied and old editions leak into training.
Research-level and guess-proof, but scores are pre-correction while the maintainer reviews about a third of the problems.
AIME on the fresh 2026 exam is the honest signal. Ignore any AIME 2024 number, which is inflated by contamination, and watch for scores quietly run with a calculator or code tool.
AIMEFrontierMath is the only test at this difficulty, but read it as provisional while the error review and the funding conflict are unresolved.
FrontierMathMath is where contamination bites hardest. A model can look brilliant on an old exam it memorized and ordinary on a fresh one, so always ask which exam a number came from. At the very top, the leaders are usually a fraction of a point apart, which is a tie, not a ranking.
The benchmark creates the need to know. The catalog explains the ideas behind it:
AIME leaders are separated by a fraction of a point, well inside the margin of error, so read the top as a cluster, not a ranking. Scores on older AIME papers run 10 to 20 points higher because those exams leaked into training; this pick uses the fresh 2026 exam.
FrontierMath numbers are contested: Epoch AI found fatal errors in roughly one third of the problems, so every current score is pre-correction, and OpenAI funded the benchmark with access to most problems. Read the standings as provisional.