Sign in
Benchmarks
Multimodal
Near ceiling(near tie)

Top models score so close together that the gaps are inside the measurement noise, so this test no longer tells you which model is better.

MMMU-Pro

MMMU-Pro is the harder, de-saturated version of MMMU: the same college-level questions that mix images with text, but with more answer choices and a vision-only mode, so a model cannot lean on the text alone or guess its way through.

What it measures

It is built to keep separating top multimodal models after they maxed out the original MMMU.

Example task

A question shows a chart, diagram, or figure and asks a college-exam-level question with up to ten answer options, sometimes with the question embedded in the image itself, so the model must actually read and reason about the visual, not just the surrounding words.

Why you should care

As AI assistants increasingly handle screenshots, documents, and photos, MMMU-Pro is the sterner test of whether a model truly understands images at a college level. Usefully, unlike the original MMMU (whose board is self-reported), this one has an independent evaluator, so the numbers are more trustworthy.

How scoring works

Accuracy on the harder multiple-choice set (higher is better). The figures here come from Artificial Analysis, an independent group that runs its own evaluations rather than taking lab submissions. The underlying benchmark is from 2024.

How to read the numbers

MMMU-Pro runs roughly 10 to 20 points lower than the original MMMU because it is deliberately harder; the current leaders cluster in the low-to-mid 80s, close enough that the ordering at the very top is within normal measurement noise.

What to watch

The top is a near-tie: the same model at two effort settings holds the first two spots within a couple of points, and the independent board's current top three are all Google Gemini variants, so it reads as a single-lab cluster rather than a broad field. Independent boards and self-reported aggregators also disagree on the leader and even the scale (one lists a 94 percent self-reported figure versus the mid-80s on independent runs), so prefer the independently-evaluated number.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders as of JUN 2026
Gemini 3.5 Flash (high)Google
84%Artificial Analysis, independent
Gemini 3.5 Flash (medium)Google
84%same model, lower effort; tied
Gemini 3.1 Pro PreviewGoogle
82%within measurement noise

The top ranks are a statistical tie. Read this as a cluster, not a clean number one.

StatusNear ceiling
Board shows no public last-updated date.
View the leaderboard