{"slug": "which-ai-model-actually-writes-best-boards-disagree", "title": "Which AI Model Actually Writes Best? Boards Disagree", "summary": "Two major AI writing leaderboards, EQ-Bench Creative Writing v3 and arena.ai's Creative Writing category, produce inverted rankings for the same models, with claude-fable-5 ranked #1 on arena.ai but 6th on EQ-Bench, and kimi-k3 ranked 2nd on EQ-Bench but #23 on arena.ai. The boards differ in methodology—arena.ai uses human blind pairwise voting while EQ-Bench uses LLM judges with a rubric and Elo—and even disagree on a factual detail, as arena.ai labels glm-5.3-max as MIT-licensed while the model card says otherwise. The findings suggest that no single leaderboard should be treated as definitive for creative writing quality.", "body_md": "If you have just read a “best AI for writing” article, the honest answer to whether you should believe it is: not on the strength of either leaderboard alone. EQ-Bench Creative Writing v3 and arena.ai’s Creative Writing category — whose page states an August 27, 2026 snapshot — do not merely re-rank the same models. They invert. claude-fable-5 is first on arena.ai and sixth on EQ-Bench; kimi-k3 is second on EQ-Bench and twenty-third on arena.ai.\n\nThe disagreement is measurable, and — more usefully — explainable. One board is a human blind-voting arena; the other is an LLM-judged rubric-plus-Elo pipeline. The full side-by-side table is directly below, and the rest of the page explains why two carefully built instruments pointed at the same task return opposite answers.\n\n- 01The boards invert, not just disagree.claude-fable-5: #1 on arena.ai, 6th on EQ-Bench, 183 Elo behind claude-opus-5. kimi-k3: 2nd on EQ-Bench, #23 on arena.ai. claude-opus-5: 1st on EQ-Bench, 10th on arena.ai.\n- 02They are different instruments, on different scales.arena.ai aggregates human blind pairwise votes; EQ-Bench uses two Claude judges for a rubric score and a pairwise Elo. The scales are anchored differently and can never be averaged or converted.\n- 03They agree on exactly one thing.Each board puts the highest-ranked open-weight model shown here just behind its #1 — 38 points on arena.ai, 45.3 Elo on EQ-Bench — two separate observations on two incompatible scales.\n- 04They even disagree on a checkable fact.arena.ai labels glm-5.3-max as MIT-licensed. The zai-org/GLM-5.3 model card says license: other with a bespoke licence name. A leaderboard’s metadata columns deserve the same skepticism as its scores.\n\n## 01 — The findingSame models, opposite *answers*.\n\nTwo definitions before the table, because both carry weight in it. **Creative Writing** is the narrow category both boards publish: fiction, storytelling, and poetry written to a prompt. Leaderboard platforms track it separately from general “Writing” tasks like emails, essays, and editing, so a rank here is a creative-writing rank, nothing broader. An **Elo score**, as leaderboards use the term, is a relative rating inferred from many pairwise comparisons: winning against a strong opponent moves a model up more than winning against a weak one. It is meaningful only within one board’s pool, solver, and anchor points — which is why the two Elo columns below run on visibly different number lines and must be read as two separate rankings, never one.\n\n| Model family | arena.ai rank · score | EQ-Bench rank · Elo | Places moved |\n|---|---|---|---|\n| claude-fable-5same name on both boards | #1 · 1505±9 | 6th · 1933.2 | Down 5 places, board to board |\n| claude-opus-5arena.ai lists claude-opus-5-high | #10 · 1471±9 | 1st · 2116.1 | Up 9 places — #1 on the other board |\n| kimi-k3arena.ai lists kimi-k3-max | #23 · 1458±11 | 2nd · 2070.8 | Up 21 places — the sharpest inversion |\n| GLM-5.3arena.ai lists glm-5.3-max, possibly a hosted-only tier | #13 · 1467±18 | 3rd · 2062.4 | Up 10 places |\n| claude-opus-4-7arena.ai lists claude-opus-4-7-high | #4 · 1489±7 | 8th · 1906.4 | Down 4 places |\n| muse-sparkarena.ai carries no version suffix; EQ-Bench lists muse-spark-1.1 — a family match, not an exact one | #20 · 1464±14 | 7th · 1915.2 | Up 13 places |\n\nRead the first three rows again. The model arena.ai’s voters rank first sits 183 Elo behind the EQ-Bench leader on EQ-Bench’s own scale. The model EQ-Bench ranks second is twenty-third with human voters. The model EQ-Bench ranks first does not crack arena.ai’s top nine. This is not noise around a shared consensus — the two boards return structurally different orderings of the same vendor lineups on the same nominal task. Neither board is *the* writing leaderboard, and any article that cites exactly one of them as settled truth has made a choice it probably has not told you about.\n\nTwo boards that agree on almost no individual ranking each place their top open-weight row just behind their top closed one — 38 points on one scale, 45.3 Elo on the other. Two separate readings that happen to point the same way.Digital Applied analysis, August 31, 2026\n\n## 02 — MethodWhat was collected, and how the boards were *matched*.\n\nThis is a comparison of two published datasets, not a benchmark we ran. That makes the matching rules — which names were paired, which scales were kept apart, which board stamps its own snapshot — the part a citing reader needs most.\n\nBoth boards were pulled first-party; no figure comes from an aggregator or a secondary write-up.\n\n- What was collected\n- Each board’s published table, read from the board itself: EQ-Bench’s rank, Elo, rubric, slop and length columns; arena.ai’s rank, score with confidence interval, vote counts, and vendor and licence labels. Plus each board’s own methodology documentation. Rows reproduced on this page are a selection.\n- Collected\n- 2026-09-01, in a single pass per board. This page is dated August 31, 2026; the figures are as collected, one day later.\n- Board-stated dates\n- arena.ai states Aug 27, 2026 on the category page, with 1,214,472 votes across 393 models in this category. EQ-Bench publishes no as-of stamp anywhere on its Creative Writing v3 page; its only dated signal is a March 1, 2026 changelog note about a judge switch.\n- Sources\n- The\n[EQ-Bench Creative Writing v3 leaderboard](https://eqbench.com/creative_writing.html)and its[methodology page](https://eqbench.com/about.html#creative-writing-v3); the[arena.ai Creative Writing leaderboard](https://arena.ai/leaderboard/text/creative-writing)(lmarena.ai now redirects there); LMSYS’s first-party[Style Control post](https://www.lmsys.org/blog/2024-08-28-style-control/); and one peer-reviewed study of LLM-judge bias,[“Judging the Judges” (TMLR 2026)](https://arxiv.org/abs/2604.23178). - Units and scales\n- Both boards publish an Elo-style score, on different scales: arena.ai’s Creative Writing scores cluster around 1500, EQ-Bench’s top out above 2100 with fixed anchors (DeepSeek-R1 at 1500, ministral-3b at 200). The two are never averaged, converted, or combined into one ranking anywhere on this page.\n- Name matching\n- arena.ai lists tier suffixes (claude-opus-5-high, kimi-k3-max, glm-5.3-max); EQ-Bench lists bare model names. Pairings in the inversion table are deliberate family-level matches, flagged per-row. muse-spark carries no version suffix on arena.ai and was matched to EQ-Bench’s muse-spark-1.1 as a family, not an exact deployment.\n- Exclusions and limitations\n- Models present on only one board are excluded from the side-by-side table. The numeric effect of arena.ai’s Style Control toggle on this category’s ranking could not be captured; only the toggle’s existence and documented design are reported. Our own post archive is not used as a comparison dataset.\n\n## 03 — The dataThe two boards, *side by side*.\n\nFirst, EQ-Bench Creative Writing v3’s top ten. Note the rubric column alongside the Elo — the same board publishes two scores per model, and section 05 is about the rows where they tell different stories. Length is the average output length the board reports per model, and it does real analytical work here.\n\n| Rank | Model | Elo | Rubric | Slop | Length (chars) |\n|---|---|---|---|---|---|\n| 1 | claude-opus-5 | 2116.1 | 85.35 | 0.9 | 6,003 |\n| 2 | kimi-k3 | 2070.8 | 84.25 | 1.3 | 5,488 |\n| 3 | GLM-5.3 | 2062.4 | 85.20 | 1.2 | 5,913 |\n| 4 | gpt-5.6-sol | 1964.1 | 83.90 | 1.6 | 8,548 |\n| 5 | ox-alpha | 1959.7 | 84.45 | 1.4 | 6,112 |\n| 6 | claude-fable-5 | 1933.2 | 84.05 | 1.4 | 5,887 |\n| 7 | muse-spark-1.1 | 1915.2 | 82.70 | 1.7 | 7,551 |\n| 8 | claude-opus-4-7 | 1906.4 | 82.85 | 1.5 | 5,692 |\n| 9 | gpt-5.6-terra | 1850.0 | 82.80 | 1.7 | 10,271 |\n| 10 | gpt-5.5 | 1844.0 | 85.05 | 1.8 | 12,945 |\n\nNow the arena.ai rows for the same families, from a page that states its own snapshot date — Aug 27, 2026 — and its own sample: 1,214,472 human votes across 393 models in this category. The confidence intervals matter: glm-5.3-max, qwen3.8-max, muse-spark, and kimi-k3-max sit within a few points of one another, with overlapping intervals, on 1,286 to 3,208 votes each.\n\n| Rank | Model | Score | Votes | Vendor · licence label |\n|---|---|---|---|---|\n| 1 | claude-fable-5 | 1505±9 | 5,160 | Anthropic |\n| 2 | claude-opus-4-6-high | 1500±7 | 12,682 | Anthropic |\n| 4 | claude-opus-4-7-high | 1489±7 | 10,548 | Anthropic |\n| 5 | gemini-3-pro | 1483±8 | 6,244 | Google · Proprietary |\n| 10 | claude-opus-5-high | 1471±9 | 6,810 | Anthropic |\n| 13 | glm-5.3-max | 1467±18 | 1,286 | Z.ai · “MIT” — disputed by the model card; see below |\n| 16 | qwen3.8-max | 1464±14 | 2,244 | Alibaba · Proprietary |\n| 20 | muse-spark | 1464±14 | 1,950 | Meta |\n| 23 | kimi-k3-max | 1458±11 | 3,208 | Moonshot · Kimi K3 licence |\n\nOne row needs a correction the board has not made. arena.ai’s licence label reads “Z.ai · MIT” for glm-5.3-max. The Hugging Face model card for [zai-org/GLM-5.3](https://huggingface.co/zai-org/GLM-5.3) states `license: other`\n\nwith `license_name: glm-5.3`\n\n— a bespoke licence, not MIT. We documented that licence in detail in [our August 28 post on the GLM-5.3 weights](/blog/glm-5-3-weights-bespoke-license-not-mit). To be precise about what was checked: the correction rests on the zai-org/GLM-5.3 model card, and glm-5.3-max may be a hosted-only tier — but a filter column that lets a reader sort by “MIT” and returns this model is disagreeing with the model’s own card. The two boards, in other words, diverge on a checkable fact as well as on scores.\n\n## 04 — The instrumentsTwo instruments, *not two opinions*.\n\nThe inversion stops being mysterious once you look at how each number is produced. arena.ai is **human pairwise voting**: two anonymous outputs shown blind, side by side, one vote for the better one, aggregated through a Bradley-Terry model — the standard statistical method for turning many pairwise win-loss records into one rating per model. EQ-Bench is **LLM-judged**, twice over: outputs are first graded against a scoring rubric by one Claude judge, then run through pairwise matchups against neighbouring models — judged by a second, different Claude judge, with win margins feeding a modified Glicko/Trueskill-style solver — a rating system of the same broad family as Elo, extended to weigh how decisively each matchup was won — run until ranks stabilise. Per the site’s March 1, 2026 note, the Elo judge is Claude Sonnet 4.6 while the rubric judge remains Claude Sonnet 4: two judges, two columns, one leaderboard.\n\n| Dimension | EQ-Bench Creative Writing v3 | arena.ai Creative Writing |\n|---|---|---|\n| Rater | Two LLM judges: Claude Sonnet 4 for the rubric, Claude Sonnet 4.6 for pairwise Elo (per the site’s March 1, 2026 note) | Anonymous humans voting blind on side-by-side outputs |\n| Unit of scoring | A rubric score per output, plus pairwise matchups with graded win margins feeding a Glicko/Trueskill-style Elo solver | One vote per pairwise comparison, aggregated by a Bradley-Terry model into a score with a confidence interval |\n| Length handling | Pairwise judging truncates outputs to a standardised length; rubric scoring keeps the full output — two policies on one board, by design | A Style Control toggle models length as an independent variable in the Bradley-Terry regression |\n| Style handling | A user-set Vocab Control slider penalising “overly complex vocab usage”, and a word-frequency Slop score built on a\n|\n\nTwo design details deserve their own sentences. First, EQ-Bench is adversarial by design — its prompts are chosen to be “challenging for weaker models and therefore highly discriminative”, and the site says plainly that “the purpose of the evaluation is not to help models write their best. Instead, we are deliberately exposing weaknesses.” arena.ai measures preference under whatever its voters bring. An instrument built to expose weaknesses and an instrument built to aggregate preference are answering different questions even when the category name matches. Second, arena.ai’s Style Control exists precisely because human votes carry style bias: LMSYS’s own launch post explains, “We explicitly model style as an independent variable in our Bradley-Terry regression. For example, we added length as a feature — just like each model, the length difference has its own Arena Score!” The toggle is live on the Creative Writing page today; what its activation does numerically to the 2026 gaps in this post could not be captured, so this page claims no narrowing or widening from it.\n\nA 2026 peer-reviewed study in Transactions on Machine Learning Research, [“Judging the Judges”](https://arxiv.org/abs/2604.23178), found style bias is the dominant LLM-judge bias — 0.10 to 0.76 across judge models, favouring markdown over plain prose, against at most 0.04 for position bias. It also found length preference is not universal: Claude-family judges in the study preferred *concise* answers (−0.12). Both of EQ-Bench’s judges are Claude models — which cuts against any simple “LLM judges always reward length” reading of the board, and makes the divergences in the next section more interesting, not less.\n\n## 05 — The nested disagreementRubric and Elo disagree inside *one board*, too.\n\nYou do not need two leaderboards to see the disagreement — EQ-Bench publishes it in two columns of its own table. gpt-5.5 scores 85.05 on the rubric against claude-opus-5’s 85.35 — a 0.3 gap on graded quality — while sitting 272 Elo lower in the pairwise ranking, at 2.2× the output length. Further down the same table, horizon-beta posts a rubric of 83.30 — level with claude-opus-4-8 — at 14,202 characters, with an Elo of 1624.5. The board publishes both columns and they do not tell the same story.\n\n| Model | Rubric | Elo | Length (chars) |\n|---|---|---|---|\n| claude-opus-5 | 85.35 | 2116.1 | 6,003 |\n| gpt-5.5 | 85.05 | 1844.0 | 12,945 |\n| horizon-beta | 83.30 | 1624.5 | 14,202 |\n\nEQ-Bench does not hide this — its methodology page addresses it directly: “Pairwise matchups allow the judge to be more discriminative than scoring a single item in isolation... The scores may also differ because we use different criteria in the judging prompts between rubric & pairwise. The judge will also be subject to different biases depending on the evaluation method.” And on which column to believe: “Why do these scores disagree? ... Which one is right? Well, both and neither.”\n\nOne more pattern from the same table, because it breaks a default assumption buyers carry: newer is not better at writing. On EQ-Bench, claude-opus-4-8 (1835.3) ranks below claude-opus-4-7 (1906.4), and claude-sonnet-5 (1787.6) ranks below claude-sonnet-4-6 (1804.5) — same-vendor version regressions, on this benchmark, published in the vendor-neutral place vendors do not control. If a leaderboard only ever confirmed release-note ordering, it would not be measuring anything.\n\n## 06 — The convergenceThe one thing both boards *agree* on.\n\nHere is the finding that survives everything above. The two boards disagree about nearly every individual rank — and still agree about the shape of the market: on each board, measured on its own scale, the highest-ranked open-weight model in the rows above sits a few dozen points behind that board’s #1 — 38 points on arena.ai, 45.3 Elo on EQ-Bench. These are two separate observations on two incompatible scales, and they concern different models. That is what makes the agreement striking rather than circular.\n\n##### Top open-weight model shown, behind #1\n\nglm-5.3-max at 1467 sits 38 points behind claude-fable-5's 1505 on the human-vote board, as stated Aug 27, 2026 — the highest-ranked open-weight family among the rows shown above.\n\n##### Top open-weight model shown, behind #1\n\nkimi-k3 at 2070.8 sits 45.3 Elo behind claude-opus-5's 2116.1 on the LLM-judged board. A different open model, a different scale, the same shape.\n\n##### Separate observations, never one number\n\nThe boards' Elo scales are anchored differently — one clusters around 1500, the other tops out above 2100. The two gaps cannot be averaged, converted, or combined into a single open-vs-closed figure.\n\nNeither board states this finding — each publishes only its own ranking, so the convergence only becomes visible when the two are laid side by side. It is also the most defensible takeaway for a buyer: whichever instrument you trust, the gap between the top open-weight models on these boards and the frontier is real but narrow, and the models occupying it are cheap to try.\n\n## 07 — ImplicationsHow to act when the boards *cancel out*.\n\nSo which board should you act on? Neither, on its own — and this post will not hand you a winner, because a single verdict is exactly the artefact the data argues against. We have published a [head-to-head frontier-model verdict](/blog/grok-4-5-vs-opus-4-8-vs-gpt-5-5-best-frontier-model-2026) where the evidence supported one; here the evidence supports a refusal. What the two boards give you instead is a pair of directional signals with known, documented biases — and EQ-Bench itself says the quiet part on its own methodology page: “The scores and rankings should only ever be interpreted as a rough guide of writing ability”, and “it’s good to be skeptical of benchmark numbers by default.”\n\n##### Pick the board whose rater resembles your audience\n\narena.ai aggregates untrained human preference under blind comparison — closer to how a newsletter subscriber or a client experiences copy. EQ-Bench is a graded, adversarial examination by Claude judges. If your writing is judged by people skimming, the human-vote signal is the nearer proxy; if it is judged on craft under scrutiny, the rubric-and-Elo signal is.\n\n##### One board quoted alone is a choice, not a fact\n\nAny article declaring a best writing model from one leaderboard has silently picked an instrument whose biases it probably has not read. Ask which board, which category, and what the other board says about the same model before repeating the claim.\n\n##### Use the one convergent finding, then run your own briefs\n\nBoth boards place a top-ranked open-weight model close behind their own #1, on their own scales. That makes a two-tier shortlist cheap: one frontier model, one open model, your actual prompts, your judgement on the outputs. Ten of your own briefs beat either board for your use case.\n\nFor reading any leaderboard — contamination, category design, cherry-picking — our [benchmark methodology guide](/blog/llm-benchmark-methodology-2026-contamination-leaderboard-guide) covers the general literacy this post applies to one specific disagreement. If the model is only half the equation for you — the other half being prompts, style constraints, and editorial scaffolding — the [free skills that measurably improve AI writing](/blog/free-skills-that-improve-ai-writing) matter more than a five-place rank difference ever will. And if the real question is standing up an editorial pipeline where model choice, briefs, and quality control are designed together rather than argued from leaderboards, that is what our [Content Engine service](/services/content-engine) builds.\n\n## 08 — ConclusionThe disagreement is the *finding*.\n\n### Neither board alone, both boards together, and one number they agree on.\n\nThese two creative-writing leaderboards invert each other on the same nominal task: first place on one is sixth on the other, second on one is twenty-third on the other. That is not a scandal and not noise — it is what happens when a human blind-voting arena and a two-judge LLM pipeline, each with documented and partly self-admitted biases, are pointed at something as contested as writing quality.\n\nWhat survives is precise: the inversion itself, the rubric-vs-Elo split inside EQ-Bench, and the one convergent finding — a top open-weight model sitting close behind #1 on each board, on two scales that must never be merged. A reader who carries those three things can evaluate any “best AI for writing” headline in about ten seconds.\n\nThe practical move is unchanged from the data: shortlist one frontier and one open model, run your own briefs, and trust your reading of the outputs over either board’s ordering. Both boards, to their credit, would tell you the same.", "url": "https://wpnews.pro/news/which-ai-model-actually-writes-best-boards-disagree", "canonical_source": "https://www.digitalapplied.com/blog/which-ai-model-writes-best-leaderboards-disagree", "published_at": "2026-08-31 00:00:00+00:00", "updated_at": "2026-09-01 08:56:09.070843+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["EQ-Bench", "arena.ai", "claude-fable-5", "kimi-k3", "claude-opus-5", "GLM-5.3", "claude-opus-4-7", "muse-spark"], "alternates": {"html": "https://wpnews.pro/news/which-ai-model-actually-writes-best-boards-disagree", "markdown": "https://wpnews.pro/news/which-ai-model-actually-writes-best-boards-disagree.md", "text": "https://wpnews.pro/news/which-ai-model-actually-writes-best-boards-disagree.txt", "jsonld": "https://wpnews.pro/news/which-ai-model-actually-writes-best-boards-disagree.jsonld"}}