Top models score so close together that the gaps are inside the measurement noise, so this test no longer tells you which model is better.
MMLU-Pro
MMLU-Pro asks exam-grade questions across 14 fields, but with 10 answer choices instead of the usual 4, so lucky guessing barely helps.
It exists because the original MMLU got too easy: top models had piled up near 90 percent, so the test could no longer tell them apart. Models score roughly 16 to 33 points lower on MMLU-Pro than on the old MMLU.
A representative item reads like a hard university final: a multi-step chemistry, law, or economics question with 10 plausible options, where ruling out the obvious wrong answers still leaves several traps that take real domain reasoning to eliminate.
When you read that a model scores 90 percent on a knowledge test, MMLU-Pro is the reason to ask which test: the original MMLU is so saturated the number is meaningless, while the same models drop sharply on the harder version. It is a clean lesson in how a benchmark stops measuring anything once everyone aces it.
Overall accuracy, the percent of about 12,000 questions answered correctly, ranked on the official TIGER-Lab leaderboard (the team that built the benchmark). The board mixes entries it evaluated itself with scores labs submit themselves, and it ranks by the single Overall column.
Random guessing scores about 10 percent because each question has 10 options. The original MMLU was already saturated near 90 percent, and on MMLU-Pro the leaders now cluster around 89 to 91 percent.
Two cautions. First, the official board was last updated in March 2026, so the newest frontier models may simply be missing. Second, only the top entry was independently evaluated by TIGER-Lab; the second and third places are scores the submitting labs reported themselves, with no subject-level breakdown, so do not read all three as equally checked. As the leaders crowd into a one-to-two point band near 90 percent, MMLU-Pro is itself starting to saturate, the same fate that retired the original MMLU.
The benchmark creates the need to know. The catalog explains the ideas behind it: