Sign in
Back to tasks
Use case

What people actually prefer

You want the model real users prefer when they cannot see the brand, a reality check on exam scores rather than a fixed-test reward.

Bottom line

On LMArena, the blind human-preference board, the top three are within about six Elo points with overlapping intervals, so there is no real number one, and all three are Anthropic models. As of JUN 2026 the listed leader is also frozen, suspended from public access, so read the top as a single-lab tie, not a live winner.

The tests that matter

Human preference is read mainly through LMArena, the blind side-by-side vote, with EQ-Bench as a more specific signal for writing quality.

LMArena (Chatbot Arena)

Millions of blind human votes, but the top three are a single-lab statistical tie and the leader is suspended from public access.

Also central to Writing and communication.

EQ-Bench (Creative Writing)

A more specific signal for one thing people vote on, writing quality, judged by a neutral third party.

Also central to Writing and communication.

How to choose
A broad sense of which model people like

LMArena is the best popularity signal, but read the top as a tie and remember it rewards what is pleasant in a short chat, which is not the same as what is correct.

LMArena (Chatbot Arena)
Preference specifically for writing

For writing quality, a judged board like EQ-Bench is more specific than a general-preference vote; see the writing use case.

EQ-Bench (Creative Writing)
What to watch

Preference boards measure what feels good in a short blind chat, which is easy to over-read: a near-tie at the top is common, single-lab clusters happen, and scores move week to week, so any snapshot ages fast. Use it as a reality check on exam scores, not as a ranking of correctness.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Best for overall human preference
claude-fable-5Anthropic
Disputed
1508 +/- 9top three span only 6 Elo
JUN 2026·arena.ai
claude-opus-4-6-thinkingAnthropic
1504 +/- 4tied with #1 within error bars
JUN 2026·arena.ai
claude-opus-4-7-thinkingAnthropic
1502 +/- 5all-Anthropic top three
JUN 2026·arena.ai
Near-tie

The top three span only about six Elo points with overlapping confidence intervals, so there is no real number one, and all three are Anthropic models, a single-lab cluster rather than a broad field. The listed leader is also frozen, suspended from public access in June 2026, so it can no longer gather new votes.