The leaderboard is current and actively updated, so this ranking reflects roughly where things stand now.
EQ-Bench (Creative Writing)
EQ-Bench Creative Writing scores how good a model is at actual creative writing, where the hard part is not producing grammatical sentences but avoiding cliche, repetition, and the bland AI slop that creeps in over a longer piece.
An independent judge model reads each sample and rates it. It is one of the few writing-quality signals not run by the labs themselves.
A model is given a creative prompt, say a short story in a specific voice, and its output is scored on craft: originality, style, how much stock phrasing or slop it leans on, and whether it repeats itself or degrades as it goes.
For the single most common student use of AI, writing and editing, this is the closest thing to a trustworthy quality score, because it is judged by a neutral third party (Sam Paech's EQ-Bench) rather than reported by the model's maker. It is a better guide than a vendor's marketing claim when you are choosing a model to help you write.
Models are ranked by an Elo Score from head-to-head comparisons (higher is better, like a chess rating), with a separate 0 to 100 Rubric Score also reported. Samples are graded by an LLM judge. Maintained independently at eqbench.com.
Elo is relative, so the number only means something next to other models: the current top sits around 2190, third place around 2020, and weaker models fall well below. There is no fixed ceiling; the score says who beats whom, not an absolute quality percentage.
Two cautions. The top two are a near-tie (about 12 Elo points apart), so do not read number one as a clear winner. And the two metrics disagree: GPT-5.5 has the highest Rubric Score of the top three yet ranks only third on Elo, so which model wins depends on which column you read. An LLM judge can also share blind spots with the models it grades, which is an open question for any LLM-judged board.
The benchmark creates the need to know. The catalog explains the ideas behind it:
The top ranks are a statistical tie. Read this as a cluster, not a clean number one.