There is no trustworthy current leaderboard, so the most recent reliable figure is already months old.
Global-MMLU
Global-MMLU asks MMLU-style exam questions, general knowledge and reasoning, but across 42 languages, and it separates questions that need culture-specific knowledge from ones that do not.
The hard part it exposes is that a model can ace a topic in English yet fail the same question in a lower-resourced language or when local cultural knowledge is required.
The same multiple-choice exam question is posed in many languages, from high-resource ones like Spanish and Chinese to lower-resourced ones like Yoruba or Sinhala, and the model is scored per language, so an English-only strength is revealed as a weakness elsewhere.
For anyone studying or working outside English, a single headline MMLU score hides large per-language gaps, and this benchmark is built to surface them. If you need a model for coursework, translation, or research in another language, the average tells you little; the per-language picture is what matters.
Accuracy on multilingual multiple-choice questions, reported per language and split into culturally sensitive versus culturally agnostic subsets. Built by Cohere Labs (Cohere For AI); published at ACL 2025. It is a dataset for running your own evaluation, not a maintained public leaderboard.
About 28 percent of the questions require culturally sensitive knowledge, and a model's ranking can change depending on whether you score the full set or only those questions. Frontier models tend to score high on high-resource languages and drop sharply on low-resource ones; there is no single headline number that captures both.
There is no public Global-MMLU leaderboard at all: the primary sources (the dataset card and the paper) ship the data for you to run yourself but publish no ranking, so any Global-MMLU top model you see online is a third-party aggregator's self-reported figure. The benchmark's own authors also flag a geographic skew in the source material (heavily North-American and European), which is part of the bias story it was built to expose.
The benchmark creates the need to know. The catalog explains the ideas behind it:
No reliable current leaderboard. The most cited figure is a dated anchor:
No public leaderboard exists: Global-MMLU is a 42-language dataset meant for running your own evaluation, and its primary sources publish no ranking or top-model scores. The durable finding (ACL 2025) is structural, not a single number: about 28 percent of questions need culturally sensitive knowledge, and model rankings shift depending on whether you score the full set or only the culturally sensitive subset.