A writers' leaderboard, not a coder's Surge AI published Hemingway-bench on 29 September, an AI writing leaderboard built from more than 5,000 blind pairwise comparisons judged by professional screenwriters, poets, speechwriters and copyeditors, which placed Google's Gemini 3 Flash and Gemini 3 Pro and Anthropic's Claude Opus 4.5 in the top three spots. The Surge AI team said the benchmark was created in part because EQ-Bench's autograder agreed with expert writers' verdicts on the same pairs as little as 43% of the time, with Gemini 3 Flash winning unanimously among human judges on a Rita Mae Brown-voice short-story prompt where EQ-Bench ranked a different model second overall. The ten-model shortlist also covered Qwen3, GPT-5.2 Chat, Kimi K2, Grok, Llama 4 Maverick and Nova, with the judges finding the best model depends on the writing task. A new AI writing benchmark, judged by working writers rather than by an autograder, has put Google’s Gemini 3 Flash and Gemini 3 Pro and Anthropic’s Claude Opus 4.5 at the top of its table. Hemingway-bench https://surgehq.ai/blog/hemingway-bench-ai-writing-leaderboard , published on 29 September by Surge AI, was built from more than 5,000 blind pairwise comparisons run by professional screenwriters, poets, speechwriters and copyeditors. The goal, the team said, was to score the qualities a reader actually cares about: voice, restraint, originality and the ability to follow a brief. What the judges found The full shortlist covered ten named models. Closed flagships took the first three places, with Gemini 3 Flash described as a master wordsmith that loved creative constraints, Gemini 3 Pro building rich worlds with specific detail afternoon light in a 1950s Philippine kitchen, the way a grandmother made adobo , and Opus 4.5, which the judges described as having the most human-like voice of the field. The runners-up, in the judges’ own personas, were where the picture got useful: - Qwen3 — the most original model on the list, but the most prone to factual errors and forced humour. A high-risk, high-reward pick for creative work. - GPT-5.2 Chat — solid for practical everyday tasks, from coordination emails to friendly advice, but not flashy on creative briefs. - Kimi K2 — competent on business writing, strained on creative. One judge called it anexecutive assistant who had studied creative writing long ago . - Grok — informal and trope-heavy; worked for casual emails, not much else. - Llama 4 Maverick — the most verbose model on the list and the one most given to placeholder phrasing on creative prompts. - Nova — structure in place, but a deeper spark missing. Why the autograder got it wrong The Surge team built Hemingway-bench in part because they had watched EQ-Bench’s autograder get reward-hacked. In their analysis, the automated grader’s verdicts aligned with expert writers’ verdicts on the same pairs as little as 43% of the time, the Hemingway-bench team said. The worked example they published was a short-story prompt that asked for a Rita Mae Brown voice. The model EQ-Bench ranked second overall, the judges wrote, leaned so hard on metaphor that the prose became incoherent; the model that EQ-Bench scored lower, Gemini 3 Flash, won unanimously among the human judges. 43%of the time, EQ-Bench’s autograder agreed with expert writers on which model wrote better. The wider argument is that automated rubrics reward the surface features of good writing — many literary devices, varied sentence structure, dense vocabulary — while readers respond to the cumulative effect. Hemingway-bench scores each response on a handful of sub-dimensions full list in the tech box and then takes a holistic view on top. How to use this for your own writing The clearest read of the Hemingway-bench results is that the right model depends on the job, and the runner-up often beats the leader for the thing you actually need to write today. - Creative writing — short stories, scripts, poetry. Gemini 3 Flash or Opus 4.5, with Qwen3 as a free local option if you can stomach the factual-error risk and have the hardware to run it see the tech box Qwen 3.8 27B is ready to download https://www.runagentrun.co.uk/articles/qwen-3-8-27b-is-ready-to-download/ for heavier briefs . - Business and professional writing — emails, marketing copy, internal docs. GPT-5.2 Chat, with Kimi K2 a credible open alternative. - Heartfelt, voice-led pieces — speeches, personal notes, obituaries. Opus 4.5, with Gemini 3 Pro a strong second. - Long-form and translation-heavy work — models with longer-context windows earn their keep on documents the smaller ones cannot hold specific token budgets in the tech box . For a worked example of the open-weights end, see our Llama 4 Scout piece https://www.runagentrun.co.uk/articles/llama-4-scout-long-context-open/ . MindStudio’s text-generation guide https://www.mindstudio.ai/blog/choosing-text-generation-model-bd422 is a useful general primer if you want to map these strengths across more use cases. Try the leaderboard prompts yourself this afternoon — a bedtime story, a thank-you email to a school, a Rita Mae Brown pastiche — and grade the output with your own rubric. The model you keep coming back to is the one for you; the one that scores well on a checklist is a different question, and one the Hemingway-bench team is asking for a reason. For a practical stack reference, our free AI tiers guide https://www.runagentrun.co.uk/articles/free-ai-tiers-got-good/ and our LLM tier roundup https://www.runagentrun.co.uk/articles/the-llm-tier-that-actually-fits-your-work/ still hold: match the model to the task, and don’t pay flagship prices for a thank-you note. Sources & quotes Every quotation in this article is verbatim from a named source — click any