Sign in
Benchmarks
Multimodal
Near ceiling(near tie)

Top models score so close together that the gaps are inside the measurement noise, so this test no longer tells you which model is better.

MMMU

MMMU tests whether a model can answer college-level exam questions that mix images with text: reading a chart, a circuit diagram, a medical scan, or a music score and reasoning about it.

What it measures

It spans six broad fields and dozens of subjects, with 30 different image types.

Example task

A question might show a chemistry diagram, an economics chart, or a sheet of music and ask a college-exam-level question that only makes sense if the model actually understands the image, not just the words around it.

Why you should care

MMMU is the main yardstick for multimodal models, the ones that see as well as read, which is the direction consumer AI is heading as assistants start handling screenshots, documents, and photos. It matters because the leaders now sit just below the human-expert reference, so multimodal understanding at a college level is largely a solved demo, though with a big asterisk about how the scores are collected.

How scoring works

Accuracy on the validation split (about 900 questions), where the official leaderboard accepts results labs submit themselves rather than running its own evaluations. That self-submission model is important context for the numbers below.

How to read the numbers

Random guessing is about 25 percent on the multiple-choice format; the Human Expert (High) reference is 88.6 percent and is excluded from the ranking. The leaders now sit in the mid-80s, just under that human reference.

What to watch

Every score on the MMMU board is self-reported: the leaderboard accepts entries labs submit themselves rather than independently re-running them, so the ordering rewards whoever reports, not necessarily whoever is best. The top scores sit within about a point of each other (a near-tie inside the roughly 1.5-point margin of error), so the ranking at the top is not meaningful. The more reliable test split has gone quiet, reportedly after its answers were released, so the validation split is now the main comparison surface. Different aggregators also crown different leaders.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders as of APR 2026
Qwen3.6 PlusAlibaba (Qwen)
Self-reported
86%self-reported; human expert ref 88.6%
APR 2026·llm-stats.com
GPT-5.1OpenAI
Self-reported
85.4%Thinking and Instant variants tie here
NOV 2025·codesota.com
GPT-5OpenAI
Self-reported
84.2%self-reported; inside the margin
AUG 2025·llm-stats.com

The top ranks are a statistical tie. Read this as a cluster, not a clean number one.

StatusNear ceiling
Board last updated JUN 2026, may be behind.
View the leaderboard