Images, charts, and documents
You want a model that understands what it is looking at, a figure, a table, a form, and reasons about it, which is where consumer AI is heading as assistants handle screenshots and photos.
As of JUN 2026 on MMMU-Pro, the independently-run test of college-level image reasoning, the top two slots are the same model at two effort settings, both at 84 percent, and the top three are all Google Gemini variants. Read the lead as a single-lab cluster measured by one evaluator, not a settled ranking across the field.
Image and document understanding is read through MMMU-Pro and MMMU for college-level visual reasoning, and OmniDocBench for turning real pages into structured data.
The harder, de-saturated version with an independent evaluator, but the top is a within-model tie and all-Google.
The main multimodal yardstick, but every score is self-reported and the leaders are within a point.
The most comprehensive document-parsing test, but the top scores are self-reported and the number one shares a maintainer with the benchmark.
MMMU-Pro is the sterner, independently-run test, so prefer it over the original MMMU, but read the top as a single-lab cluster and judge on the images you actually care about.
MMMU-ProOmniDocBench is the most complete test, but its numbers are self-reported with a maintainer conflict, so run your own pages before trusting a leader.
OmniDocBenchMultimodal boards are unusually self-reported, and the current leaders are heavily concentrated in one vendor, so a high number can reflect who submitted, or who measured, as much as who is best. Prefer the independently-run board and cite it by name, since different evaluators rank different models on top.
The benchmark creates the need to know. The catalog explains the ideas behind it:
The top two slots are the same model, Gemini 3.5 Flash, at two effort settings, not two different models, and the top three are all Google Gemini variants. These are Artificial Analysis numbers, which can differ from the official MMMU-Pro harness, so treat them as one lab's measurement and cite that board by name.
Every OmniDocBench score is self-reported in the project's own repository rather than independently re-run, and the number one is maintained by OpenDataLab, the same organization that maintains the benchmark, a direct conflict of interest. There is no independent board to check it against.