Languages and translation
You want a model that holds its quality outside English, across coursework, translation, and research in lower-resourced languages, not one that only shines in English.
No board here measures this cleanly, so we are not naming a leader. The honest finding is structural, not a ranking.
Multilingual ability is read through Global-MMLU, the one test built to expose per-language gaps, but it ships as a dataset with no public ranking.
Built to expose per-language gaps across 42 languages, but it is a dataset for running your own evaluation, with no public leaderboard at all.
No board here measures this cleanly, so we are not naming a winner. Judge it yourself on these things.
Global-MMLU is the right framework, but there is no ranking to quote.
No clean answer yet: Its primary sources publish the dataset but no leaderboard, so any Global-MMLU top model you see online is a third-party aggregator's self-reported figure. Test the model in your own language.
- Does it stay accurate in your specific language, or only in English, when you ask the same question both ways?
- Does it handle culturally specific knowledge in your region, not just a translated North-American or European default?
A single headline knowledge score hides large per-language gaps: a model can ace a topic in English and fail the same question in a lower-resourced language. There is no public multilingual ranking to lean on, and the dataset's authors flag a geographic skew in the source material, so test in your own language rather than trusting an aggregate.
The benchmark creates the need to know. The catalog explains the ideas behind it:
No public leaderboard exists. The durable, dated finding (ACL 2025) is structural: about 28 percent of Global-MMLU questions need culturally sensitive knowledge, and model rankings shift depending on whether you score the full set or only that subset.
No public Global-MMLU leaderboard exists at all: the primary sources ship the dataset for you to run yourself but publish no ranking, so there is no trustworthy current top model to name.