Sign in
Back to tasks
Use case

Factual accuracy: not making things up

You want a model that gets short, checkable facts right and declines the ones it does not know, instead of inventing a confident, wrong answer.

Bottom line

On the Kaggle SimpleQA board, Gemini 3.1 Pro Preview tops it at about 74.8 percent as of JUN 2026, clear of second place beyond the error bars. But the name SimpleQA is overloaded across several different boards, so this figure only means something when you say it came from this specific board.

The tests that matter

Factual honesty is read through SimpleQA for short checkable facts, with Humanity's Last Exam as a cross-check on confident wrong answers at the expert frontier.

SimpleQA

A current, regularly-updated board for short factual recall, but provenance is mixed and the name is shared by several different boards.

Humanity's Last Exam

A cross-check at the expert frontier: it penalizes confident wrong answers, even though it is not a pure factual-recall test.

Also central to Science and expert knowledge.

How to choose
Short, single-answer factual questions

The Kaggle SimpleQA board is the live signal, but name it specifically, since the DeepMind and llm-stats boards report very different numbers for tests with the same name.

SimpleQA
Whether a model admits what it does not know

Prefer the F-score view where available, since it rewards declining over guessing wrong; the Correct column shown here only counts right answers and does not punish a confident miss as much.

SimpleQA
What to watch

Hallucination is the failure mode that most often burns students, and SimpleQA is the best-known test of it, but it is also a lesson in naming: several different boards all called SimpleQA report wildly different numbers, and a multi-model average sits near 21 percent, so a very high figure usually means web search or a generous board, not raw knowledge. Always name the board.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

SimpleQA
As of JUN 2026
Best for short, checkable factual recall
Gemini 3.1 Pro PreviewGoogle
74.8% plus or minus 1.4%Correct column; 2024 frontier was high-40s
JUN 2026ยทkaggle.com
Caveat

The top three are all Google Gemini variants, so the lead rests on one vendor's standing on one board, and provenance is mixed, Kaggle reproduces some scores but others can be publisher-reported. The name SimpleQA is overloaded, so always cite this specific Kaggle OpenAI board.