The headline numbers are disputed, self-reported, or under revision, so treat the ranking as a claim, not a settled fact.
FrontierMath
FrontierMath is a set of original research-level math problems so hard they take specialist mathematicians hours or days to solve, and at launch in late 2024 every leading model scored below 2 percent.
The problems are guess-proof (under a 1 percent chance of a lucky correct answer) and checked automatically, spanning everything from number theory to algebraic geometry.
A FrontierMath problem looks less like a school exam and more like a research exercise: a precisely stated question whose answer is a specific number or object, requiring the kind of work a math PhD might spend an afternoon on. The hardest tier (Tier 4) is genuine research-mathematician territory.
FrontierMath is the cleanest test of whether AI can do real, novel mathematical reasoning rather than recite memorized results, and scores have rocketed from under 2 percent to the high 80s on the hardest tier in about eighteen months. It also comes with a built-in lesson about trusting benchmark numbers: the benchmark was quietly funded by OpenAI, which had privileged access to the problems, and its maintainer recently found errors in roughly a third of the questions.
Accuracy, the percent of problems solved, with answers auto-verified by exact match or symbolic computation. It is run by Epoch AI, which evaluates models independently (often with pre-release access). Scores are reported by tier, where Tier 4 is the hardest, research-level band; the figures here are from the LM Council aggregator, which draws on Epoch's evaluations.
All frontier models scored below 2 percent at the November 2024 launch; expert mathematicians sampled on the easier tiers averaged about 19 percent. Today's leaders reach the high 80s on the hardest tier, a jump that is part of why the benchmark is being re-examined.
FrontierMath carries two disclosed problems that make its numbers contested. First, Epoch AI announced in May 2026 that an AI-assisted review found fatal errors in roughly one third of the problems, so corrected scores are pending and every current figure is pre-correction. Second, OpenAI funded the benchmark and had access to most problems and solutions, a conflict Epoch disclosed after the fact; the often-quoted 25 percent o3 result from December 2024 was self-reported from an internal high-compute run, and Epoch's own test of the public model found closer to 10 percent. Read the standings as provisional.
The benchmark creates the need to know. The catalog explains the ideas behind it: