cd /news/artificial-intelligence/which-ai-model-actually-writes-best-… · home topics artificial-intelligence article
[ARTICLE · art-117521] src=digitalapplied.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Which AI Model Actually Writes Best? Boards Disagree

Two major AI writing leaderboards, EQ-Bench Creative Writing v3 and arena.ai's Creative Writing category, produce inverted rankings for the same models, with claude-fable-5 ranked #1 on arena.ai but 6th on EQ-Bench, and kimi-k3 ranked 2nd on EQ-Bench but #23 on arena.ai. The boards differ in methodology—arena.ai uses human blind pairwise voting while EQ-Bench uses LLM judges with a rubric and Elo—and even disagree on a factual detail, as arena.ai labels glm-5.3-max as MIT-licensed while the model card says otherwise. The findings suggest that no single leaderboard should be treated as definitive for creative writing quality.

read16 min views1 publishedAug 31, 2026
Which AI Model Actually Writes Best? Boards Disagree
Image: Digitalapplied (auto-discovered)

If you have just read a “best AI for writing” article, the honest answer to whether you should believe it is: not on the strength of either leaderboard alone. EQ-Bench Creative Writing v3 and arena.ai’s Creative Writing category — whose page states an August 27, 2026 snapshot — do not merely re-rank the same models. They invert. claude-fable-5 is first on arena.ai and sixth on EQ-Bench; kimi-k3 is second on EQ-Bench and twenty-third on arena.ai. The disagreement is measurable, and — more usefully — explainable. One board is a human blind-voting arena; the other is an LLM-judged rubric-plus-Elo pipeline. The full side-by-side table is directly below, and the rest of the page explains why two carefully built instruments pointed at the same task return opposite answers.

  • 01The boards invert, not just disagree.claude-fable-5: #1 on arena.ai, 6th on EQ-Bench, 183 Elo behind claude-opus-5. kimi-k3: 2nd on EQ-Bench, #23 on arena.ai. claude-opus-5: 1st on EQ-Bench, 10th on arena.ai.
  • 02They are different instruments, on different scales.arena.ai aggregates human blind pairwise votes; EQ-Bench uses two Claude judges for a rubric score and a pairwise Elo. The scales are anchored differently and can never be averaged or converted.
  • 03They agree on exactly one thing.Each board puts the highest-ranked open-weight model shown here just behind its #1 — 38 points on arena.ai, 45.3 Elo on EQ-Bench — two separate observations on two incompatible scales.
  • 04They even disagree on a checkable fact.arena.ai labels glm-5.3-max as MIT-licensed. The zai-org/GLM-5.3 model card says license: other with a bespoke licence name. A leaderboard’s metadata columns deserve the same skepticism as its scores.

01 — The findingSame models, opposite answers. #

Two definitions before the table, because both carry weight in it. Creative Writing is the narrow category both boards publish: fiction, storytelling, and poetry written to a prompt. Leaderboard platforms track it separately from general “Writing” tasks like emails, essays, and editing, so a rank here is a creative-writing rank, nothing broader. An Elo score, as leaderboards use the term, is a relative rating inferred from many pairwise comparisons: winning against a strong opponent moves a model up more than winning against a weak one. It is meaningful only within one board’s pool, solver, and anchor points — which is why the two Elo columns below run on visibly different number lines and must be read as two separate rankings, never one.

Model family arena.ai rank · score EQ-Bench rank · Elo Places moved
claude-fable-5same name on both boards #1 · 1505±9 6th · 1933.2 Down 5 places, board to board
claude-opus-5arena.ai lists claude-opus-5-high #10 · 1471±9 1st · 2116.1 Up 9 places — #1 on the other board
kimi-k3arena.ai lists kimi-k3-max #23 · 1458±11 2nd · 2070.8 Up 21 places — the sharpest inversion
GLM-5.3arena.ai lists glm-5.3-max, possibly a hosted-only tier #13 · 1467±18 3rd · 2062.4 Up 10 places
claude-opus-4-7arena.ai lists claude-opus-4-7-high #4 · 1489±7 8th · 1906.4 Down 4 places
muse-sparkarena.ai carries no version suffix; EQ-Bench lists muse-spark-1.1 — a family match, not an exact one #20 · 1464±14 7th · 1915.2 Up 13 places

Read the first three rows again. The model arena.ai’s voters rank first sits 183 Elo behind the EQ-Bench leader on EQ-Bench’s own scale. The model EQ-Bench ranks second is twenty-third with human voters. The model EQ-Bench ranks first does not crack arena.ai’s top nine. This is not noise around a shared consensus — the two boards return structurally different orderings of the same vendor lineups on the same nominal task. Neither board is the writing leaderboard, and any article that cites exactly one of them as settled truth has made a choice it probably has not told you about.

Two boards that agree on almost no individual ranking each place their top open-weight row just behind their top closed one — 38 points on one scale, 45.3 Elo on the other. Two separate readings that happen to point the same way.Digital Applied analysis, August 31, 2026

02 — MethodWhat was collected, and how the boards were matched. #

This is a comparison of two published datasets, not a benchmark we ran. That makes the matching rules — which names were paired, which scales were kept apart, which board stamps its own snapshot — the part a citing reader needs most.

Both boards were pulled first-party; no figure comes from an aggregator or a secondary write-up.

  • What was collected
  • Each board’s published table, read from the board itself: EQ-Bench’s rank, Elo, rubric, slop and length columns; arena.ai’s rank, score with confidence interval, vote counts, and vendor and licence labels. Plus each board’s own methodology documentation. Rows reproduced on this page are a selection.
  • Collected
  • 2026-09-01, in a single pass per board. This page is dated August 31, 2026; the figures are as collected, one day later.
  • Board-stated dates
  • arena.ai states Aug 27, 2026 on the category page, with 1,214,472 votes across 393 models in this category. EQ-Bench publishes no as-of stamp anywhere on its Creative Writing v3 page; its only dated signal is a March 1, 2026 changelog note about a judge switch.
  • Sources
  • The EQ-Bench Creative Writing v3 leaderboardand itsmethodology page; thearena.ai Creative Writing leaderboard(lmarena.ai now redirects there); LMSYS’s first-partyStyle Control post; and one peer-reviewed study of LLM-judge bias,“Judging the Judges” (TMLR 2026). - Units and scales
  • Both boards publish an Elo-style score, on different scales: arena.ai’s Creative Writing scores cluster around 1500, EQ-Bench’s top out above 2100 with fixed anchors (DeepSeek-R1 at 1500, ministral-3b at 200). The two are never averaged, converted, or combined into one ranking anywhere on this page.
  • Name matching
  • arena.ai lists tier suffixes (claude-opus-5-high, kimi-k3-max, glm-5.3-max); EQ-Bench lists bare model names. Pairings in the inversion table are deliberate family-level matches, flagged per-row. muse-spark carries no version suffix on arena.ai and was matched to EQ-Bench’s muse-spark-1.1 as a family, not an exact deployment.
  • Exclusions and limitations
  • Models present on only one board are excluded from the side-by-side table. The numeric effect of arena.ai’s Style Control toggle on this category’s ranking could not be captured; only the toggle’s existence and documented design are reported. Our own post archive is not used as a comparison dataset.

03 — The dataThe two boards, side by side. #

First, EQ-Bench Creative Writing v3’s top ten. Note the rubric column alongside the Elo — the same board publishes two scores per model, and section 05 is about the rows where they tell different stories. Length is the average output length the board reports per model, and it does real analytical work here.

Rank Model Elo Rubric Slop Length (chars)
1 claude-opus-5 2116.1 85.35 0.9 6,003
2 kimi-k3 2070.8 84.25 1.3 5,488
3 GLM-5.3 2062.4 85.20 1.2 5,913
4 gpt-5.6-sol 1964.1 83.90 1.6 8,548
5 ox-alpha 1959.7 84.45 1.4 6,112
6 claude-fable-5 1933.2 84.05 1.4 5,887
7 muse-spark-1.1 1915.2 82.70 1.7 7,551
8 claude-opus-4-7 1906.4 82.85 1.5 5,692
9 gpt-5.6-terra 1850.0 82.80 1.7 10,271
10 gpt-5.5 1844.0 85.05 1.8 12,945

Now the arena.ai rows for the same families, from a page that states its own snapshot date — Aug 27, 2026 — and its own sample: 1,214,472 human votes across 393 models in this category. The confidence intervals matter: glm-5.3-max, qwen3.8-max, muse-spark, and kimi-k3-max sit within a few points of one another, with overlapping intervals, on 1,286 to 3,208 votes each.

Rank Model Score Votes Vendor · licence label
1 claude-fable-5 1505±9 5,160 Anthropic
| 2 | claude-opus-4-6-high | 1500±7 | 12,682 | Anthropic |
| 4 | claude-opus-4-7-high | 1489±7 | 10,548 | Anthropic |

| 5 | gemini-3-pro | 1483±8 | 6,244 | Google · Proprietary | | 10 | claude-opus-5-high | 1471±9 | 6,810 | Anthropic | | 13 | glm-5.3-max | 1467±18 | 1,286 | Z.ai · “MIT” — disputed by the model card; see below | | 16 | qwen3.8-max | 1464±14 | 2,244 | Alibaba · Proprietary | | 20 | muse-spark | 1464±14 | 1,950 | Meta | | 23 | kimi-k3-max | 1458±11 | 3,208 | Moonshot · Kimi K3 licence |

One row needs a correction the board has not made. arena.ai’s licence label reads “Z.ai · MIT” for glm-5.3-max. The Hugging Face model card for zai-org/GLM-5.3 states license: other

with license_name: glm-5.3 — a bespoke licence, not MIT. We documented that licence in detail in our August 28 post on the GLM-5.3 weights. To be precise about what was checked: the correction rests on the zai-org/GLM-5.3 model card, and glm-5.3-max may be a hosted-only tier — but a filter column that lets a reader sort by “MIT” and returns this model is disagreeing with the model’s own card. The two boards, in other words, diverge on a checkable fact as well as on scores.

04 — The instrumentsTwo instruments, not two opinions. #

The inversion stops being mysterious once you look at how each number is produced. arena.ai is human pairwise voting: two anonymous outputs shown blind, side by side, one vote for the better one, aggregated through a Bradley-Terry model — the standard statistical method for turning many pairwise win-loss records into one rating per model. EQ-Bench is LLM-judged, twice over: outputs are first graded against a scoring rubric by one Claude judge, then run through pairwise matchups against neighbouring models — judged by a second, different Claude judge, with win margins feeding a modified Glicko/Trueskill-style solver — a rating system of the same broad family as Elo, extended to weigh how decisively each matchup was won — run until ranks stabilise. Per the site’s March 1, 2026 note, the Elo judge is Claude Sonnet 4.6 while the rubric judge remains Claude Sonnet 4: two judges, two columns, one leaderboard.

Dimension EQ-Bench Creative Writing v3 arena.ai Creative Writing
Rater Two LLM judges: Claude Sonnet 4 for the rubric, Claude Sonnet 4.6 for pairwise Elo (per the site’s March 1, 2026 note) Anonymous humans voting blind on side-by-side outputs
Unit of scoring A rubric score per output, plus pairwise matchups with graded win margins feeding a Glicko/Trueskill-style Elo solver One vote per pairwise comparison, aggregated by a Bradley-Terry model into a score with a confidence interval
Length handling Pairwise judging truncates outputs to a standardised length; rubric scoring keeps the full output — two policies on one board, by design A Style Control toggle models length as an independent variable in the Bradley-Terry regression
Style handling A user-set Vocab Control slider penalising “overly complex vocab usage”, and a word-frequency Slop score built on a

Two design details deserve their own sentences. First, EQ-Bench is adversarial by design — its prompts are chosen to be “challenging for weaker models and therefore highly discriminative”, and the site says plainly that “the purpose of the evaluation is not to help models write their best. Instead, we are deliberately exposing weaknesses.” arena.ai measures preference under whatever its voters bring. An instrument built to expose weaknesses and an instrument built to aggregate preference are answering different questions even when the category name matches. Second, arena.ai’s Style Control exists precisely because human votes carry style bias: LMSYS’s own launch post explains, “We explicitly model style as an independent variable in our Bradley-Terry regression. For example, we added length as a feature — just like each model, the length difference has its own Arena Score!” The toggle is live on the Creative Writing page today; what its activation does numerically to the 2026 gaps in this post could not be captured, so this page claims no narrowing or widening from it.

A 2026 peer-reviewed study in Transactions on Machine Learning Research, “Judging the Judges”, found style bias is the dominant LLM-judge bias — 0.10 to 0.76 across judge models, favouring markdown over plain prose, against at most 0.04 for position bias. It also found length preference is not universal: Claude-family judges in the study preferred concise answers (−0.12). Both of EQ-Bench’s judges are Claude models — which cuts against any simple “LLM judges always reward length” reading of the board, and makes the divergences in the next section more interesting, not less.

05 — The nested disagreementRubric and Elo disagree inside one board, too. #

You do not need two leaderboards to see the disagreement — EQ-Bench publishes it in two columns of its own table. gpt-5.5 scores 85.05 on the rubric against claude-opus-5’s 85.35 — a 0.3 gap on graded quality — while sitting 272 Elo lower in the pairwise ranking, at 2.2× the output length. Further down the same table, horizon-beta posts a rubric of 83.30 — level with claude-opus-4-8 — at 14,202 characters, with an Elo of 1624.5. The board publishes both columns and they do not tell the same story.

Model Rubric Elo Length (chars)
claude-opus-5 85.35 2116.1 6,003
gpt-5.5 85.05 1844.0 12,945
horizon-beta 83.30 1624.5 14,202

EQ-Bench does not hide this — its methodology page addresses it directly: “Pairwise matchups allow the judge to be more discriminative than scoring a single item in isolation... The scores may also differ because we use different criteria in the judging prompts between rubric & pairwise. The judge will also be subject to different biases depending on the evaluation method.” And on which column to believe: “Why do these scores disagree? ... Which one is right? Well, both and neither.”

One more pattern from the same table, because it breaks a default assumption buyers carry: newer is not better at writing. On EQ-Bench, claude-opus-4-8 (1835.3) ranks below claude-opus-4-7 (1906.4), and claude-sonnet-5 (1787.6) ranks below claude-sonnet-4-6 (1804.5) — same-vendor version regressions, on this benchmark, published in the vendor-neutral place vendors do not control. If a leaderboard only ever confirmed release-note ordering, it would not be measuring anything.

06 — The convergenceThe one thing both boards agree on. #

Here is the finding that survives everything above. The two boards disagree about nearly every individual rank — and still agree about the shape of the market: on each board, measured on its own scale, the highest-ranked open-weight model in the rows above sits a few dozen points behind that board’s #1 — 38 points on arena.ai, 45.3 Elo on EQ-Bench. These are two separate observations on two incompatible scales, and they concern different models. That is what makes the agreement striking rather than circular.

Top open-weight model shown, behind #1

glm-5.3-max at 1467 sits 38 points behind claude-fable-5's 1505 on the human-vote board, as stated Aug 27, 2026 — the highest-ranked open-weight family among the rows shown above.

Top open-weight model shown, behind #1

kimi-k3 at 2070.8 sits 45.3 Elo behind claude-opus-5's 2116.1 on the LLM-judged board. A different open model, a different scale, the same shape.

Separate observations, never one number

The boards' Elo scales are anchored differently — one clusters around 1500, the other tops out above 2100. The two gaps cannot be averaged, converted, or combined into a single open-vs-closed figure.

Neither board states this finding — each publishes only its own ranking, so the convergence only becomes visible when the two are laid side by side. It is also the most defensible takeaway for a buyer: whichever instrument you trust, the gap between the top open-weight models on these boards and the frontier is real but narrow, and the models occupying it are cheap to try.

07 — ImplicationsHow to act when the boards cancel out. #

So which board should you act on? Neither, on its own — and this post will not hand you a winner, because a single verdict is exactly the artefact the data argues against. We have published a head-to-head frontier-model verdict where the evidence supported one; here the evidence supports a refusal. What the two boards give you instead is a pair of directional signals with known, documented biases — and EQ-Bench itself says the quiet part on its own methodology page: “The scores and rankings should only ever be interpreted as a rough guide of writing ability”, and “it’s good to be skeptical of benchmark numbers by default.”

Pick the board whose rater resembles your audience

arena.ai aggregates untrained human preference under blind comparison — closer to how a newsletter subscriber or a client experiences copy. EQ-Bench is a graded, adversarial examination by Claude judges. If your writing is judged by people skimming, the human-vote signal is the nearer proxy; if it is judged on craft under scrutiny, the rubric-and-Elo signal is.

One board quoted alone is a choice, not a fact

Any article declaring a best writing model from one leaderboard has silently picked an instrument whose biases it probably has not read. Ask which board, which category, and what the other board says about the same model before repeating the claim.

Use the one convergent finding, then run your own briefs

Both boards place a top-ranked open-weight model close behind their own #1, on their own scales. That makes a two-tier shortlist cheap: one frontier model, one open model, your actual prompts, your judgement on the outputs. Ten of your own briefs beat either board for your use case.

For reading any leaderboard — contamination, category design, cherry-picking — our benchmark methodology guide covers the general literacy this post applies to one specific disagreement. If the model is only half the equation for you — the other half being prompts, style constraints, and editorial scaffolding — the free skills that measurably improve AI writing matter more than a five-place rank difference ever will. And if the real question is standing up an editorial pipeline where model choice, briefs, and quality control are designed together rather than argued from leaderboards, that is what our Content Engine service builds.

08 — ConclusionThe disagreement is the finding. #

Neither board alone, both boards together, and one number they agree on.

These two creative-writing leaderboards invert each other on the same nominal task: first place on one is sixth on the other, second on one is twenty-third on the other. That is not a scandal and not noise — it is what happens when a human blind-voting arena and a two-judge LLM pipeline, each with documented and partly self-admitted biases, are pointed at something as contested as writing quality.

What survives is precise: the inversion itself, the rubric-vs-Elo split inside EQ-Bench, and the one convergent finding — a top open-weight model sitting close behind #1 on each board, on two scales that must never be merged. A reader who carries those three things can evaluate any “best AI for writing” headline in about ten seconds.

The practical move is unchanged from the data: shortlist one frontier and one open model, run your own briefs, and trust your reading of the outputs over either board’s ordering. Both boards, to their credit, would tell you the same.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @eq-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/which-ai-model-actua…] indexed:0 read:16min 2026-08-31 ·