Top models score so close together that the gaps are inside the measurement noise, so this test no longer tells you which model is better.
MMMU-Pro
MMMU-Pro is the harder, de-saturated version of MMMU: the same college-level questions that mix images with text, but with more answer choices and a vision-only mode, so a model cannot lean on the text alone or guess its way through.
It is built to keep separating top multimodal models after they maxed out the original MMMU.
A question shows a chart, diagram, or figure and asks a college-exam-level question with up to ten answer options, sometimes with the question embedded in the image itself, so the model must actually read and reason about the visual, not just the surrounding words.
As AI assistants increasingly handle screenshots, documents, and photos, MMMU-Pro is the sterner test of whether a model truly understands images at a college level. Usefully, unlike the original MMMU (whose board is self-reported), this one has an independent evaluator, so the numbers are more trustworthy.
Accuracy on the harder multiple-choice set (higher is better). The figures here come from Artificial Analysis, an independent group that runs its own evaluations rather than taking lab submissions. The underlying benchmark is from 2024.
MMMU-Pro runs roughly 10 to 20 points lower than the original MMMU because it is deliberately harder; the current leaders cluster in the low-to-mid 80s, close enough that the ordering at the very top is within normal measurement noise.
The top is a near-tie: the same model at two effort settings holds the first two spots within a couple of points, and the independent board's current top three are all Google Gemini variants, so it reads as a single-lab cluster rather than a broad field. Independent boards and self-reported aggregators also disagree on the leader and even the scale (one lists a 94 percent self-reported figure versus the mid-80s on independent runs), so prefer the independently-evaluated number.
The benchmark creates the need to know. The catalog explains the ideas behind it:
The top ranks are a statistical tie. Read this as a cluster, not a clean number one.