Science and expert knowledge
You want a model that holds up on hard, specialist questions, the kind where a confident wrong answer is worse than useful, across biology, physics, chemistry, and beyond.
On Humanity's Last Exam, the test built to be the hardest exam in the world for an AI, the top two models are a genuine statistical tie, their intervals overlap, and even the leaders score under 50 percent. As of MAY 2026 read the top as shared, and remember a meaningful share of the answer key is itself disputed.
Expert knowledge is read through GPQA Diamond for graduate science, Humanity's Last Exam for the frontier of human knowledge, and MMLU-Pro for broad exam-grade reasoning.
Google-proof graduate science, but saturated: the leaders sit within about a point, so the ordering at the top is close to a coin flip.
Still genuinely unsolved at under 50 percent, but the top is a tie and reviewers found many reference answers wrong.
Also central to Factual accuracy: not making things up.
Broad exam-grade knowledge, but the official board froze in March 2026 and is itself starting to saturate near 90 percent.
GPQA Diamond is the standard, but it is saturated, so treat the top handful as interchangeable and pick on access and cost.
GPQA DiamondHumanity's Last Exam still has room to measure, which makes it the best progress gauge, but read the top as a tie and the absolute numbers as approximate given the answer-key issues.
Humanity's Last ExamMMLU-Pro covers the widest range, but its board is months old and only the top entry was independently evaluated, so newer models may simply be missing.
MMLU-ProTwo traps here. Most of these tests are saturating, so a one or two point gap at the top is noise, not a winner. And a high score is only as trustworthy as the answer key behind it, which on Humanity's Last Exam is itself partly disputed. Prefer the tests that still have room to separate models.
The benchmark creates the need to know. The catalog explains the ideas behind it:
The top two share rank one: their confidence intervals overlap, so naming a single winner overstates the certainty. And read every absolute score with care, independent reviewers estimate 18 to 29 percent of the chemistry and biology reference answers are wrong.
GPQA Diamond is widely described as saturated: the top models cluster within about one point, smaller than normal run-to-run variation, and different independent evaluators crown different number ones. The ordering at the very top is not meaningful.