The leaderboard is current and actively updated, so this ranking reflects roughly where things stand now.
Humanity's Last Exam
Humanity's Last Exam is a 2,500-question test built to be the hardest exam in the world for an AI, spanning math, the humanities, and the natural sciences at the absolute frontier of human knowledge.
It was designed so even the best models would fail most of it: the top score today is still under 50 percent.
Questions are written by experts to stump frontier models, mixing deep specialist knowledge with multi-step reasoning, and about one in seven requires reading a diagram or figure. A typical item is the kind of question only a working researcher in that exact subfield could answer from memory.
When labs say their model is approaching the limits of human knowledge, this is the test they mean. It matters because it is one of the few benchmarks the best models still fail badly, so it actually has room to measure progress, unlike the saturated tests where everyone scores in the 90s. It is also a cautionary tale: independent reviewers found a meaningful share of its answer key is wrong.
Accuracy on the 2,500-question set, run independently by Scale AI (the benchmark was built with the Center for AI Safety). The board uses a Rank (Upper Bound) method, so models whose confidence intervals overlap share a rank instead of being forced into a false order.
Random guessing is near zero because most answers are open-ended; the test is designed so even human experts would clear it only within their own specialties. The prior state of the art before mid-2025 frontier models was below 10 percent; the current leaders sit in the mid-40s.
Two big cautions. First, the top of the board is a genuine statistical tie: the leading two models' confidence intervals overlap, so naming a single winner overstates the certainty. Second, and more serious, independent analyses estimate that 18 to 29 percent of the chemistry and biology reference answers are wrong (one widely cited example marked a synthetic element that existed for milliseconds as the rarest noble gas on Earth), and a separate review flagged over a thousand items needing revision. Every score here is measured against that uncorrected answer key, so read the absolute numbers with care. A corrected version, HLE-Rolling, has been announced.
The benchmark creates the need to know. The catalog explains the ideas behind it:
The top ranks are a statistical tie. Read this as a cluster, not a clean number one.