Sign in
Back to tasks
Use case

Images, charts, and documents

You want a model that understands what it is looking at, a figure, a table, a form, and reasons about it, which is where consumer AI is heading as assistants handle screenshots and photos.

Bottom line

As of JUN 2026 on MMMU-Pro, the independently-run test of college-level image reasoning, the top two slots are the same model at two effort settings, both at 84 percent, and the top three are all Google Gemini variants. Read the lead as a single-lab cluster measured by one evaluator, not a settled ranking across the field.

The tests that matter

Image and document understanding is read through MMMU-Pro and MMMU for college-level visual reasoning, and OmniDocBench for turning real pages into structured data.

MMMU-Pro

The harder, de-saturated version with an independent evaluator, but the top is a within-model tie and all-Google.

MMMU

The main multimodal yardstick, but every score is self-reported and the leaders are within a point.

OmniDocBench

The most comprehensive document-parsing test, but the top scores are self-reported and the number one shares a maintainer with the benchmark.

How to choose
Reading charts, diagrams, and figures

MMMU-Pro is the sterner, independently-run test, so prefer it over the original MMMU, but read the top as a single-lab cluster and judge on the images you actually care about.

MMMU-Pro
Turning documents and scans into data

OmniDocBench is the most complete test, but its numbers are self-reported with a maintainer conflict, so run your own pages before trusting a leader.

OmniDocBench
What to watch

Multimodal boards are unusually self-reported, and the current leaders are heavily concentrated in one vendor, so a high number can reflect who submitted, or who measured, as much as who is best. Prefer the independently-run board and cite it by name, since different evaluators rank different models on top.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

MMMU-Pro
As of JUN 2026
Best for reading images and charts at a college level
Gemini 3.5 Flash (high)Google
84%Artificial Analysis, independent
Caveat

The top two slots are the same model, Gemini 3.5 Flash, at two effort settings, not two different models, and the top three are all Google Gemini variants. These are Artificial Analysis numbers, which can differ from the official MMMU-Pro harness, so treat them as one lab's measurement and cite that board by name.

Best for turning PDFs and scans into structured data
MinerU2.5-ProOpenDataLab
Self-reportedDisputed
95.75 Overallself-reported; same org as the benchmark
JUN 2026·github.com
Caveat

Every OmniDocBench score is self-reported in the project's own repository rather than independently re-run, and the number one is maintained by OpenDataLab, the same organization that maintains the benchmark, a direct conflict of interest. There is no independent board to check it against.