Factual accuracy: not making things up
You want a model that gets short, checkable facts right and declines the ones it does not know, instead of inventing a confident, wrong answer.
On the Kaggle SimpleQA board, Gemini 3.1 Pro Preview tops it at about 74.8 percent as of JUN 2026, clear of second place beyond the error bars. But the name SimpleQA is overloaded across several different boards, so this figure only means something when you say it came from this specific board.
Factual honesty is read through SimpleQA for short checkable facts, with Humanity's Last Exam as a cross-check on confident wrong answers at the expert frontier.
A current, regularly-updated board for short factual recall, but provenance is mixed and the name is shared by several different boards.
A cross-check at the expert frontier: it penalizes confident wrong answers, even though it is not a pure factual-recall test.
Also central to Science and expert knowledge.
The Kaggle SimpleQA board is the live signal, but name it specifically, since the DeepMind and llm-stats boards report very different numbers for tests with the same name.
SimpleQAPrefer the F-score view where available, since it rewards declining over guessing wrong; the Correct column shown here only counts right answers and does not punish a confident miss as much.
SimpleQAHallucination is the failure mode that most often burns students, and SimpleQA is the best-known test of it, but it is also a lesson in naming: several different boards all called SimpleQA report wildly different numbers, and a multi-model average sits near 21 percent, so a very high figure usually means web search or a generous board, not raw knowledge. Always name the board.
The benchmark creates the need to know. The catalog explains the ideas behind it:
The top three are all Google Gemini variants, so the lead rests on one vendor's standing on one board, and provenance is mixed, Kaggle reproduces some scores but others can be publisher-reported. The name SimpleQA is overloaded, so always cite this specific Kaggle OpenAI board.