Sign in
Benchmarks
Reasoning
Near ceiling(near tie)

Top models score so close together that the gaps are inside the measurement noise, so this test no longer tells you which model is better.

GPQA Diamond

GPQA Diamond is a set of 198 science questions so hard that PhD-level domain experts answer only about two thirds correctly, and skilled non-experts score around 34 percent even with the whole internet to help.

What it measures

The questions are deliberately Google-proof: you cannot just look them up. It is the go-to test for graduate-level reasoning in biology, physics, and chemistry.

Example task

A typical question is a graduate-level physics or chemistry problem with four answer choices, written so that searching the web turns up plausible but wrong leads. Even a specialist with unlimited time and access lands around two thirds correct.

Why you should care

GPQA Diamond is the test people point to when they say AI now reasons at a graduate level in the sciences. It matters because the questions are built so you cannot Google your way out, so a high score is harder to fake than on an open-book trivia test. The catch is that the top models have nearly run out of room at the top.

How scoring works

Accuracy, the percent of the 198 four-choice questions answered correctly. There is no official maintainer-run leaderboard; the figures here come from Artificial Analysis, an independent group that runs its own evaluations rather than accepting lab submissions, so these are independent measurements, not self-reports.

How to read the numbers

Random chance is 25 percent on these four-choice questions; PhD domain experts score roughly 65 to 70 percent; the prior state of the art before mid-2025 frontier models was around 78 to 80 percent. Today's leaders sit in the low-to-mid 90s, which is why the benchmark is called near saturated.

What to watch

GPQA Diamond is widely described as saturated: the leaders cluster within roughly one point of each other, a gap smaller than the normal run-to-run variation, so the ordering at the very top is close to a coin flip. Different independent evaluators (Artificial Analysis, Epoch AI, vals.ai) even crown different number-one models, which is itself a sign the race is inside the noise. Watch out for scores from gated, unreleased lab previews (one aggregator lists a model around 94.6 percent that the public cannot access), and remember no single board is canonical here.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders as of JUN 2026
Gemini 3.1 Pro PreviewGoogle DeepMind
94.1%Artificial Analysis, independent eval
GPT-5.5 (xhigh)OpenAI
93.5%0.9 points covers the whole top 3
GPT-5.5 (high)OpenAI
93.2%inside run-to-run noise

The top ranks are a statistical tie. Read this as a cluster, not a clean number one.

StatusNear ceiling
Board shows no public last-updated date.
View the leaderboard