Sign in
Benchmarks
Human preference
Contested(near tie)

The headline numbers are disputed, self-reported, or under revision, so treat the ranking as a claim, not a settled fact.

LMArena (Chatbot Arena)

LMArena ranks models by which one real people prefer in blind, side-by-side chats: you ask a question, see two anonymous answers, and vote for the better one, and millions of those votes become a single score.

What it measures

It is the closest thing to a popularity contest decided by users rather than a fixed exam.

Example task

You type a prompt and get two replies with the model names hidden, pick the one you like better, and only afterward see which models you compared. Millions of these pairwise votes are pooled into an Elo rating, the same math used to rank chess players.

Why you should care

LMArena is the leaderboard that tries to measure what people actually prefer, not what a fixed test rewards, which is why labs cite it constantly when a new model launches. It matters as a reality check on exam scores, but it is easy to over-read: as of mid-2026 the top three are a statistical tie, they are all from one lab, and the number-one model has been pulled from public access.

How scoring works

An Elo rating from blind pairwise human votes, shown with a confidence interval (for example, 1508 plus or minus 9). It is run by LMArena, the project formerly known as Chatbot Arena, now at arena.ai. Higher is better and there is no ceiling; 1000 is the baseline.

How to read the numbers

Elo starts at 1000 for a baseline model; GPT-4 at launch was around 1350, and the prior frontier before 2025 sat around 1250 to 1300. Today's leaders are near 1500, but the top three are within about 6 points of each other, which is inside their error bars.

What to watch

This is a genuine both-at-once case: a near-tie and a contested leader. The top three span only about 6 Elo points with overlapping confidence intervals, so there is no real number one, and all three are Anthropic models, so the top of the board is a single-lab cluster rather than a broad field. The current leader, Claude Fable 5, is flagged disputed because it was suspended from public access in June 2026 under a US export-control directive and can no longer gather new votes; it stays on the board as a frozen historical entry. Arena scores also move week to week, so any snapshot ages quickly.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Leaders as of JUN 2026
claude-fable-5Anthropic
Disputed
1508 +/- 9top three span only 6 Elo
JUN 2026·arena.ai
claude-opus-4-6-thinkingAnthropic
1504 +/- 4tied with #1 within error bars
JUN 2026·arena.ai
claude-opus-4-7-thinkingAnthropic
1502 +/- 5all-Anthropic top three
JUN 2026·arena.ai

The top ranks are a statistical tie. Read this as a cluster, not a clean number one.

StatusContested
Board last updated JUN 2026, may be behind.
View the leaderboard