cd /news/large-language-models/claude-opus-5-5-tops-mazur-constrain… · home › topics › large-language-models › article
[ARTICLE · art-145022] src=runtimewire.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Claude Opus 5.5 Tops Mazur Constrained Story Writing Benchmark

Anthropic's Claude Opus 5.5 ranked first on the Mazur Constrained Story Writing 50-task sample with a mean score of 1.00 at an estimated $0.0242 per task, according to the published leaderboard. OpenAI's GPT-6 Astra and GPT-6.1 Sol Pro tied for second at 0.99, with Sol Pro costing $0.0121 per task, while Anthropic's Claude Sonnet 5.5 placed fourth at 0.98 and $0.0106 per task. The 50 open-ended tasks were graded 0–10 by OpenAI's GPT-5.6 Sol Pro against a fixed rubric and normalized to a 0–1 scale, with reasoning effort pinned to max for every model exposing such a control.

by read2 min views2 publishedOct 4, 2026
Claude Opus 5.5 Tops Mazur Constrained Story Writing Benchmark
Image: Runtimewire (auto-discovered)

Anthropic’s Claude Opus 5.5 scored 1.00 on the 50-task sample at $0.0242 per task. OpenAI’s GPT-6 Astra and GPT-6.1 Sol Pro tied for second at 0.99, with Sol Pro costing $0.0121 per task.

        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
        · Published 

Leaderboard

| Rank | Model | Mean score | Est. cost/task |

|---|---|---|---|
| 1 | Anthropic: [Claude Opus 5.5](https://runtimewire.com/models/anthropic/claude-opus-5.5) | 1.00 | $0.0242 | 
| 2 | OpenAI: [GPT-6 Astra](https://runtimewire.com/models/native-openai/gpt-6-astra-576a599a64cb3dde) | 0.99 | $0.0599 | 
| 3 | OpenAI: [GPT-6.1 Sol](https://runtimewire.com/models/native-openai/gpt-6.1-sol-dedfefd51120f283) Pro | 0.99 | $0.0121 | 

| 4 | Anthropic: Claude Sonnet 5.5 | 0.98 | $0.0106 | | 5 | Z.ai: GLM 5.3 Prime | 0.97 | $0.0111 | | 6 | xAI: Grok Latest | 0.84 | $0.0077 | | 7 | Google: Gemini 3.8 Flash | 0.70 | $0.0047 | | 8 | Qwen: Qwen3.8 Max Prime | 0.66 | $0.0141 | | 9 | StepFun: Step 3.7 Flash | 0.49 | $0.0014 |

How we scored it

Every model answered the same 50-task battery from Mazur Constrained Story Writing (50-task sample) (Lech Mazur — LLM Creative Story-Writing Benchmark; published constrained briefs and archived V4 grading protocol), one task at a time, with no tools and no retries on content.

50 open-ended tasks were graded 0–10 by OpenAI: GPT-5.6 Sol Pro against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.

Reasoning effort was pinned to max for every model whose lane exposes a control (OpenAI-style reasoning_effort, Anthropic extended thinking, Gemini thinking config). Models marked vendor default on the interactive board expose no such control and ran as shipped.

A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.

Grading could not be completed for every task on 9 models; their means cover the graded subset.

Explore every prompt, answer, and per-task grade in the interactive leaderboard.

── more in #large-language-models 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-opus-5-5-tops…] indexed:0 read:2min 2026-10-04 · —