Sign in
Benchmarks
Reasoning
Live(near tie)

The leaderboard is current and actively updated, so this ranking reflects roughly where things stand now.

Humanity's Last Exam

Humanity's Last Exam is a 2,500-question test built to be the hardest exam in the world for an AI, spanning math, the humanities, and the natural sciences at the absolute frontier of human knowledge.

What it measures

It was designed so even the best models would fail most of it: the top score today is still under 50 percent.

Example task

Questions are written by experts to stump frontier models, mixing deep specialist knowledge with multi-step reasoning, and about one in seven requires reading a diagram or figure. A typical item is the kind of question only a working researcher in that exact subfield could answer from memory.

Why you should care

When labs say their model is approaching the limits of human knowledge, this is the test they mean. It matters because it is one of the few benchmarks the best models still fail badly, so it actually has room to measure progress, unlike the saturated tests where everyone scores in the 90s. It is also a cautionary tale: independent reviewers found a meaningful share of its answer key is wrong.

How scoring works

Accuracy on the 2,500-question set, run independently by Scale AI (the benchmark was built with the Center for AI Safety). The board uses a Rank (Upper Bound) method, so models whose confidence intervals overlap share a rank instead of being forced into a false order.

How to read the numbers

Random guessing is near zero because most answers are open-ended; the test is designed so even human experts would clear it only within their own specialties. The prior state of the art before mid-2025 frontier models was below 10 percent; the current leaders sit in the mid-40s.

What to watch

Two big cautions. First, the top of the board is a genuine statistical tie: the leading two models' confidence intervals overlap, so naming a single winner overstates the certainty. Second, and more serious, independent analyses estimate that 18 to 29 percent of the chemistry and biology reference answers are wrong (one widely cited example marked a synthetic element that existed for milliseconds as the rarest noble gas on Earth), and a separate review flagged over a thousand items needing revision. Every score here is measured against that uncorrected answer key, so read the absolute numbers with care. A corrected version, HLE-Rolling, has been announced.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders as of MAY 2026
tied
gemini-3.1-pro-preview (thinking high)Google DeepMind
46.44 +/- 1.96shares rank 1; intervals overlap
MAY 2026·labs.scale.com
tied
gpt-5.4-pro-2026-03-05OpenAI
44.32 +/- 1.95shares rank 1 with Gemini
MAY 2026·labs.scale.com
Muse SparkMeta Superintelligence Labs
40.56 +/- 1.92Scale AI independent eval
MAY 2026·labs.scale.com

The top ranks are a statistical tie. Read this as a cluster, not a clean number one.

StatusLive
Board last updated MAY 2026, may be behind.
View the leaderboard