What people actually prefer
You want the model real users prefer when they cannot see the brand, a reality check on exam scores rather than a fixed-test reward.
On LMArena, the blind human-preference board, the top three are within about six Elo points with overlapping intervals, so there is no real number one, and all three are Anthropic models. As of JUN 2026 the listed leader is also frozen, suspended from public access, so read the top as a single-lab tie, not a live winner.
Human preference is read mainly through LMArena, the blind side-by-side vote, with EQ-Bench as a more specific signal for writing quality.
Millions of blind human votes, but the top three are a single-lab statistical tie and the leader is suspended from public access.
Also central to Writing and communication.
A more specific signal for one thing people vote on, writing quality, judged by a neutral third party.
Also central to Writing and communication.
LMArena is the best popularity signal, but read the top as a tie and remember it rewards what is pleasant in a short chat, which is not the same as what is correct.
LMArena (Chatbot Arena)For writing quality, a judged board like EQ-Bench is more specific than a general-preference vote; see the writing use case.
EQ-Bench (Creative Writing)Preference boards measure what feels good in a short blind chat, which is easy to over-read: a near-tie at the top is common, single-lab clusters happen, and scores move week to week, so any snapshot ages fast. Use it as a reality check on exam scores, not as a ranking of correctness.
The benchmark creates the need to know. The catalog explains the ideas behind it:
The top three span only about six Elo points with overlapping confidence intervals, so there is no real number one, and all three are Anthropic models, a single-lab cluster rather than a broad field. The listed leader is also frozen, suspended from public access in June 2026, so it can no longer gather new votes.