Sign in
Benchmarks
Reasoning
Near ceiling

Top models score so close together that the gaps are inside the measurement noise, so this test no longer tells you which model is better.

MMLU-Pro

MMLU-Pro asks exam-grade questions across 14 fields, but with 10 answer choices instead of the usual 4, so lucky guessing barely helps.

What it measures

It exists because the original MMLU got too easy: top models had piled up near 90 percent, so the test could no longer tell them apart. Models score roughly 16 to 33 points lower on MMLU-Pro than on the old MMLU.

Example task

A representative item reads like a hard university final: a multi-step chemistry, law, or economics question with 10 plausible options, where ruling out the obvious wrong answers still leaves several traps that take real domain reasoning to eliminate.

Why you should care

When you read that a model scores 90 percent on a knowledge test, MMLU-Pro is the reason to ask which test: the original MMLU is so saturated the number is meaningless, while the same models drop sharply on the harder version. It is a clean lesson in how a benchmark stops measuring anything once everyone aces it.

How scoring works

Overall accuracy, the percent of about 12,000 questions answered correctly, ranked on the official TIGER-Lab leaderboard (the team that built the benchmark). The board mixes entries it evaluated itself with scores labs submit themselves, and it ranks by the single Overall column.

How to read the numbers

Random guessing scores about 10 percent because each question has 10 options. The original MMLU was already saturated near 90 percent, and on MMLU-Pro the leaders now cluster around 89 to 91 percent.

What to watch

Two cautions. First, the official board was last updated in March 2026, so the newest frontier models may simply be missing. Second, only the top entry was independently evaluated by TIGER-Lab; the second and third places are scores the submitting labs reported themselves, with no subject-level breakdown, so do not read all three as equally checked. As the leaders crowd into a one-to-two point band near 90 percent, MMLU-Pro is itself starting to saturate, the same fate that retired the original MMLU.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders as of MAR 2026
Gemini-3.1-ProGoogle DeepMind
91.16% overall accuracyTIGER-Lab evaluated; random guess about 10%
MAR 2026·huggingface.co
Gemini-3-Pro(11/25)Google DeepMind
Self-reported
90.1% overall accuracyself-reported by the lab
MAR 2026·huggingface.co
GPT-o1OpenAI
Self-reported
89.3% overall accuracyself-reported; the board's own label
MAR 2026·huggingface.co
StatusNear ceiling
Board last updated MAR 2026, may be behind.
View the leaderboard