Sign in
Benchmarks
Multilingual
Dated

There is no trustworthy current leaderboard, so the most recent reliable figure is already months old.

Global-MMLU

Global-MMLU asks MMLU-style exam questions, general knowledge and reasoning, but across 42 languages, and it separates questions that need culture-specific knowledge from ones that do not.

What it measures

The hard part it exposes is that a model can ace a topic in English yet fail the same question in a lower-resourced language or when local cultural knowledge is required.

Example task

The same multiple-choice exam question is posed in many languages, from high-resource ones like Spanish and Chinese to lower-resourced ones like Yoruba or Sinhala, and the model is scored per language, so an English-only strength is revealed as a weakness elsewhere.

Why you should care

For anyone studying or working outside English, a single headline MMLU score hides large per-language gaps, and this benchmark is built to surface them. If you need a model for coursework, translation, or research in another language, the average tells you little; the per-language picture is what matters.

How scoring works

Accuracy on multilingual multiple-choice questions, reported per language and split into culturally sensitive versus culturally agnostic subsets. Built by Cohere Labs (Cohere For AI); published at ACL 2025. It is a dataset for running your own evaluation, not a maintained public leaderboard.

How to read the numbers

About 28 percent of the questions require culturally sensitive knowledge, and a model's ranking can change depending on whether you score the full set or only those questions. Frontier models tend to score high on high-resource languages and drop sharply on low-resource ones; there is no single headline number that captures both.

What to watch

There is no public Global-MMLU leaderboard at all: the primary sources (the dataset card and the paper) ship the data for you to run yourself but publish no ranking, so any Global-MMLU top model you see online is a third-party aggregator's self-reported figure. The benchmark's own authors also flag a geographic skew in the source material (heavily North-American and European), which is part of the bias story it was built to expose.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders

No reliable current leaderboard. The most cited figure is a dated anchor:

No public leaderboard exists: Global-MMLU is a 42-language dataset meant for running your own evaluation, and its primary sources publish no ranking or top-model scores. The durable finding (ACL 2025) is structural, not a single number: about 28 percent of questions need culturally sensitive knowledge, and model rankings shift depending on whether you score the full set or only the culturally sensitive subset.

StatusDated
View the board