Sign in
Back to tasks
Use case

Science and expert knowledge

You want a model that holds up on hard, specialist questions, the kind where a confident wrong answer is worse than useful, across biology, physics, chemistry, and beyond.

Bottom line

On Humanity's Last Exam, the test built to be the hardest exam in the world for an AI, the top two models are a genuine statistical tie, their intervals overlap, and even the leaders score under 50 percent. As of MAY 2026 read the top as shared, and remember a meaningful share of the answer key is itself disputed.

The tests that matter

Expert knowledge is read through GPQA Diamond for graduate science, Humanity's Last Exam for the frontier of human knowledge, and MMLU-Pro for broad exam-grade reasoning.

GPQA Diamond

Google-proof graduate science, but saturated: the leaders sit within about a point, so the ordering at the top is close to a coin flip.

Humanity's Last Exam

Still genuinely unsolved at under 50 percent, but the top is a tie and reviewers found many reference answers wrong.

Also central to Factual accuracy: not making things up.

MMLU-Pro

Broad exam-grade knowledge, but the official board froze in March 2026 and is itself starting to saturate near 90 percent.

How to choose
Hard graduate-level science questions

GPQA Diamond is the standard, but it is saturated, so treat the top handful as interchangeable and pick on access and cost.

GPQA Diamond
The genuine frontier of human knowledge

Humanity's Last Exam still has room to measure, which makes it the best progress gauge, but read the top as a tie and the absolute numbers as approximate given the answer-key issues.

Humanity's Last Exam
Broad, general exam-grade knowledge

MMLU-Pro covers the widest range, but its board is months old and only the top entry was independently evaluated, so newer models may simply be missing.

MMLU-Pro
What to watch

Two traps here. Most of these tests are saturating, so a one or two point gap at the top is noise, not a winner. And a high score is only as trustworthy as the answer key behind it, which on Humanity's Last Exam is itself partly disputed. Prefer the tests that still have room to separate models.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Humanity's Last Exam
As of MAY 2026
Best at the frontier of expert knowledge
gemini-3.1-pro-preview (thinking high)Google DeepMind
46.44 +/- 1.96shares rank 1; intervals overlap
MAY 2026·labs.scale.com
gpt-5.4-pro-2026-03-05OpenAI
44.32 +/- 1.95shares rank 1 with Gemini
MAY 2026·labs.scale.com
Near-tie

The top two share rank one: their confidence intervals overlap, so naming a single winner overstates the certainty. And read every absolute score with care, independent reviewers estimate 18 to 29 percent of the chemistry and biology reference answers are wrong.

Best for graduate-level science questions
Gemini 3.1 Pro PreviewGoogle DeepMind
94.1%Artificial Analysis, independent eval
Caveat

GPQA Diamond is widely described as saturated: the top models cluster within about one point, smaller than normal run-to-run variation, and different independent evaluators crown different number ones. The ordering at the very top is not meaningful.