Writing and communication
You want a model that writes with craft, original, varied, and readable across a longer piece, not one that pads with stock phrasing and repeats itself.
On EQ-Bench, the independent judged writing board, the top two models are about 12 Elo points apart, close enough to read as a shared lead, as of JUN 2026. And the two metrics disagree: the model with the highest rubric score ranks only third on the head-to-head Elo, so which one wins depends on which column you read.
Writing quality is read mainly through EQ-Bench Creative Writing, judged by a neutral third party, with LMArena as a broader human-preference cross-check.
A neutral, LLM-judged writing signal rather than a vendor claim, but the top two are nearly tied and the two metrics disagree.
A broader human-preference signal that includes, but is not specific to, writing.
Also central to What people actually prefer.
Read the EQ-Bench Elo, but treat the top two as a shared lead rather than a clean winner.
EQ-Bench (Creative Writing)The separate rubric score ranks models differently; if you care about clean, correct prose more than head-to-head flavor, follow the rubric column instead of the Elo.
EQ-Bench (Creative Writing)LMArena adds a general human-preference signal, but it is not writing-specific; see the what people prefer use case.
LMArena (Chatbot Arena)For the most common student use of AI, writing, a neutral judged board beats a vendor's marketing claim, but two cautions remain: the top is often a near-tie, and the two metrics here can disagree, so the winner depends on which column you read. An LLM judge may also share blind spots with the models it grades.
The benchmark creates the need to know. The catalog explains the ideas behind it:
The top two are about 12 Elo apart, close enough to read as a shared lead rather than a clean number one. An LLM judge can also share blind spots with the models it grades, an open question for any judged board.
This is a metric-split pick: on EQ-Bench's separate 0 to 100 rubric score, the model that ranks third on head-to-head Elo actually scores highest, so if you weight polished craft over head-to-head preference, the ranking flips. Read the two metrics as answering different questions.