The leaderboard is current and actively updated, so this ranking reflects roughly where things stand now.
SimpleQA
SimpleQA tests whether a model knows short, checkable facts without making things up, using 4,326 deliberately hard fact-seeking questions with single, indisputable answers.
The point is not difficulty of reasoning but honesty about knowledge: a good model answers what it knows and declines what it does not, instead of confidently inventing an answer.
A question has one correct short answer, a specific date, name, or number, chosen to be obscure enough that models often get it wrong. The reply is graded correct, incorrect, or not attempted, so confidently wrong answers are penalized and an honest I do not know beats a fabrication.
Hallucination, confidently stating false facts, is the failure mode that most often burns students using AI for research or coursework, and SimpleQA is the best-known test of it. The Gemini 3.x generation recently jumped into the 70s on the Kaggle board shown here, a real step-change from the high-40s a year earlier; the catch is that the name SimpleQA is overloaded across several different boards, so a number only means something once you say which board it came from.
Models are graded correct, incorrect, or not attempted; the figure here is the Correct column, the percent of all 4,326 questions answered correctly, not the F-score SimpleQA also reports. Built by OpenAI (October 2024). The board shown is the Kaggle SimpleQA Leaderboard run by OpenAI, last updated June 2026. Provenance is mixed: Kaggle reproduces scores where it can and captures auditable outputs, but some entries can be publisher-reported, so do not read the top number as fully independent.
SimpleQA was built to be hard. Across the 2024 to 2025 frontier, models scored in the high-40s to about 50 percent (o3 about 50.5, Claude and Grok 4 about 48 to 49). The Gemini 3.x generation jumped to the 70s on this board, a recent step-change; a separate aggregator tracking roughly 45 models still shows an average near 21 percent, so a very high figure from any single board deserves a second look.
Read the board name carefully. The bars here are the Correct column, not the F-score, and the name SimpleQA is overloaded: this Kaggle OpenAI board (74.8 and 70.5 at the top) is a different board from the DeepMind SimpleQA Verified board (figures around 77, 72, and 70) and from llm-stats.com (which uses a different 0 to 1 metric). The top three are all Google Gemini variants, so the lead rests on one vendor's standing on a single board, and the third-place score could not be independently confirmed. Always name this specific board when you quote a SimpleQA number.
The benchmark creates the need to know. The catalog explains the ideas behind it: