Sign in
Back to tasks
Use case

Writing and communication

You want a model that writes with craft, original, varied, and readable across a longer piece, not one that pads with stock phrasing and repeats itself.

Bottom line

On EQ-Bench, the independent judged writing board, the top two models are about 12 Elo points apart, close enough to read as a shared lead, as of JUN 2026. And the two metrics disagree: the model with the highest rubric score ranks only third on the head-to-head Elo, so which one wins depends on which column you read.

The tests that matter

Writing quality is read mainly through EQ-Bench Creative Writing, judged by a neutral third party, with LMArena as a broader human-preference cross-check.

EQ-Bench (Creative Writing)

A neutral, LLM-judged writing signal rather than a vendor claim, but the top two are nearly tied and the two metrics disagree.

LMArena (Chatbot Arena)

A broader human-preference signal that includes, but is not specific to, writing.

Also central to What people actually prefer.

How to choose
Head-to-head writing preference

Read the EQ-Bench Elo, but treat the top two as a shared lead rather than a clean winner.

EQ-Bench (Creative Writing)
Polished craft on a fixed rubric

The separate rubric score ranks models differently; if you care about clean, correct prose more than head-to-head flavor, follow the rubric column instead of the Elo.

EQ-Bench (Creative Writing)
A broader sense of what readers like

LMArena adds a general human-preference signal, but it is not writing-specific; see the what people prefer use case.

LMArena (Chatbot Arena)
What to watch

For the most common student use of AI, writing, a neutral judged board beats a vendor's marketing claim, but two cautions remain: the top is often a near-tie, and the two metrics here can disagree, so the winner depends on which column you read. An LLM judge may also share blind spots with the models it grades.

To go deeper

The benchmark creates the need to know. The catalog explains the ideas behind it:

Best for head-to-head writing preference
claude-fable-5Anthropic
Elo 2191.8Rubric 84.05 of 100; LLM-judged
JUN 2026·eqbench.com
claude-opus-4-7Anthropic
Elo 2179.312.5 Elo behind the leader
JUN 2026·eqbench.com
Near-tie

The top two are about 12 Elo apart, close enough to read as a shared lead rather than a clean number one. An LLM judge can also share blind spots with the models it grades, an open question for any judged board.

Best by writing-quality rubric, not head-to-head
gpt-5.5OpenAI
Elo 2019.0highest Rubric (85.05) yet third on Elo
JUN 2026·eqbench.com
Caveat

This is a metric-split pick: on EQ-Bench's separate 0 to 100 rubric score, the model that ranks third on head-to-head Elo actually scores highest, so if you weight polished craft over head-to-head preference, the ranking flips. Read the two metrics as answering different questions.