Top models score so close together that the gaps are inside the measurement noise, so this test no longer tells you which model is better.
MMMU
MMMU tests whether a model can answer college-level exam questions that mix images with text: reading a chart, a circuit diagram, a medical scan, or a music score and reasoning about it.
It spans six broad fields and dozens of subjects, with 30 different image types.
A question might show a chemistry diagram, an economics chart, or a sheet of music and ask a college-exam-level question that only makes sense if the model actually understands the image, not just the words around it.
MMMU is the main yardstick for multimodal models, the ones that see as well as read, which is the direction consumer AI is heading as assistants start handling screenshots, documents, and photos. It matters because the leaders now sit just below the human-expert reference, so multimodal understanding at a college level is largely a solved demo, though with a big asterisk about how the scores are collected.
Accuracy on the validation split (about 900 questions), where the official leaderboard accepts results labs submit themselves rather than running its own evaluations. That self-submission model is important context for the numbers below.
Random guessing is about 25 percent on the multiple-choice format; the Human Expert (High) reference is 88.6 percent and is excluded from the ranking. The leaders now sit in the mid-80s, just under that human reference.
Every score on the MMMU board is self-reported: the leaderboard accepts entries labs submit themselves rather than independently re-running them, so the ordering rewards whoever reports, not necessarily whoever is best. The top scores sit within about a point of each other (a near-tie inside the roughly 1.5-point margin of error), so the ranking at the top is not meaningful. The more reliable test split has gone quiet, reportedly after its answers were released, so the validation split is now the main comparison surface. Different aggregators also crown different leaders.
The benchmark creates the need to know. The catalog explains the ideas behind it:
The top ranks are a statistical tie. Read this as a cluster, not a clean number one.