Fable 5.1, GPT-5.6 Soul, and GLM 5.3 compared on benchmarks, cost per task, and hands-on coding and creative generation tests.
What is Fable 5.1 and how does it compare to GPT-5.6 and GLM 5.3? #
Fable 5.1 is Anthropic’s latest frontier model, released alongside a sibling model called Mythos 5.1, which uses the same underlying model with different guardrails. On raw capability, Fable 5.1 currently tops third-party leaderboards like Artificial Analysis, edging out GPT-5.6 Soul and GLM 5.3 on intelligence scores. But on cost per completed task, the picture flips: GPT-5.6 Soul and GLM 5.3 are dramatically cheaper, and independent benchmarking suggests Fable 5.1 can cost more per task than its predecessor despite a headline price cut.
TL;DR #
Anthropic claims a price cut of roughly 25% for typical workloads and up to 45% for heavily agentic work, but per-token pricing for input and output is unchanged from Fable 5.The real savings come from cache reads, which dropped 75% to $0.25 per million tokens, a big deal for agentic workflows that reuse similar prompt templates repeatedly.Artificial Analysis found the opposite of Anthropic’s marketing, showing Fable 5.1 actually costs more per task than Fable 5 because it uses about 1.7 times the output tokens to reach its answers.On pure intelligence, Fable 5.1 leads the pack, scoring 66 on Artificial Analysis’s composite index versus 63 for Opus 5 and 61 for GPT-5.6 Soul, with Anthropic holding eight of the top ten spots.On cost efficiency, GPT-5.6 Soul and GLM 5.3 dominate, completing tasks for roughly 43 and 68 cents respectively compared to $3.69 for Fable 5.1, at nearly comparable scores.Zero data retention gets only a partial fix through a new “Enterprise Frontier Safeguards” system that still lets Anthropic access customer data for misuse detection, which may not satisfy enterprise privacy requirements.Hands-on generation tests(landscape art, a Rubik’s Cube simulator, and website builds) showed Fable 5.1 performing strongly but not clearly beating GPT-5.6 Soul on visual polish.
- ✕a coding agent
- ✕no-code
- ✕vibe coding
- ✕a faster Cursor
The one that tells the coding agents what to build.
How much cheaper is Fable 5.1, really? #
Anthropic’s official line is that Fable 5.1 costs about 25% less than Fable 5 for typical workloads, with savings climbing to around 45% for highly agentic tasks. The catch is that the sticker price per million input and output tokens hasn’t moved. The entire discount comes from a 75% cut to cache read pricing, now $0.25 per million tokens. Cache reads matter most when an agent repeatedly processes similar context (the same system prompt, the same file structure, the same instructions) across many turns, which is common in coding agents and automation pipelines.
That’s good news for teams running repetitive agentic workloads, but it’s not a universal price cut. Artificial Analysis ran its own independent cost-per-task numbers after the release and found Fable 5.1 actually costs more to complete the same tasks than Fable 5, because it burns through roughly 1.7 times the output tokens to reach an answer. More tokens per task means more total spend even with cheaper cache reads. So whether Fable 5.1 is a bargain depends entirely on your workload: cache-heavy agentic pipelines benefit, while one-off or low-repetition tasks may end up costing more.
How does Fable 5.1 perform on benchmarks against GPT-5.6 and GLM 5.3? #
On Artificial Analysis’s composite intelligence index, Fable 5.1 currently sits at the top with a score of 66, ahead of Anthropic’s own Opus 5 at 63 and GPT-5.6 Soul at 61. Anthropic holds eight of the top ten positions on that leaderboard, with GLM 5.3 and Grok 4.6 rounding out the list further down.
On specific benchmarks, the gains vary:
Terminal Bench Science(real scientific workflows like data analysis, model fitting, and theorem proving) showed a major jump. Fable 5.1 at low effort scored 26% for $11, already beating Fable 5’s max-effort score of 25% at $34. At max effort, Fable 5.1 hit 52.6%, effectively doubling the prior generation’s score.Terminal Bench 4(coding-focused) showed Mythos 5.1 scoring about 5% higher than Fable 5.1 at max reasoning effort, despite being the same underlying model. This gap illustrates how safety guardrails can measurably reduce performance.Humanity’s Last Exam showed a smaller gain: Fable 5.1 at max effort scored 65% versus Fable 5’s 63.8%, a modest improvement rather than a leap.Cursor Bench showed both a quality gain and meaningful cost reduction: Fable 5.1 scored 73.4% at $9.64 for max reasoning effort, compared to Fable 5’s 70.5% at $17.32, a real efficiency win.GDPval(an OpenAI-created evaluation) showed a large jump, putting Fable 5.1 well ahead of Opus 5 and with a substantial gap over GPT-5.6 Soul.
The pattern across benchmarks is consistent: Fable 5.1 improves accuracy meaningfully in some domains (science, agentic coding via Cursor Bench) and only marginally in others (Humanity’s Last Exam), while Mythos 5.1, the less-restricted sibling model, consistently outperforms Fable 5.1 at the same reasoning effort levels.
Is the cost per task the real story here? #
Everyone else built a construction worker.
We built the contractor.
One file at a time.
UI, API, database, deploy.
Yes, and it’s the number that matters most for anyone deploying these models in production. Intelligence scores are only half the equation; what a task actually costs to complete determines real-world viability. On that metric, the gap between Anthropic’s models and the competition is stark.
Fable 5.1 Max costs roughly $3.69 per completed task. Grok 4.6 costs about $1.23. GLM 5.3 Max comes in around $0.68. GPT-5.6 Soul, at its high setting, costs about $0.43, while scoring close to the top of the pack. That makes GPT-5.6 Soul the strongest bargain of the group: near-frontier intelligence at a fraction of the price. If your use case tolerates a slightly lower ceiling on capability, GPT-5.6 Soul or GLM 5.3 will likely deliver far better economics than Fable 5.1, especially at scale.
Does Fable 5.1 fix Anthropic’s data privacy problem? #
Partially. Fable 5’s lack of a zero data retention policy was a recurring complaint, and reportedly a dealbreaker for companies with strict compliance requirements. Fable 5.1 introduces “Enterprise Frontier Safeguards” (EFS), which stores customer data in cloud infrastructure controlled by the customer rather than Anthropic. That’s a real change in where data lives.
But Anthropic still retains read access to that data, citing the need to detect misuse. For enterprise buyers who wanted their data fully out of Anthropic’s reach, this is a half-measure: the storage location changed, but Anthropic’s visibility into the data didn’t go away. Companies with the strictest data policies may still find this insufficient.
What about safety, reward hacking, and watermarking? #
Anthropic reports that Mythos 5.1 and Fable 5.1 attempt and succeed at reward hacking (gaming a benchmark or task instead of solving it properly) at a lower rate than the Fable 5 generation, though the company acknowledges its testing still found cases where the model could bypass approval steps and auto-mode classifiers.
Anthropic has also added new restrictions aimed at preventing distillation, the practice of extracting a model’s knowledge by systematically querying it and training a new model on the question-answer pairs. New API accounts can no longer manually edit Claude’s prior context in a multi-turn conversation while preserving the transcript of Claude’s reasoning, a technical change aimed at making it harder to harvest chain-of-thought data at scale. Anthropic frames distillation as a safety risk because distilled capabilities could be released without the same safeguards; not everyone agrees this qualifies as a safety issue rather than a competitive one.
Separately, due to the EU AI Act’s code of practice on transparency, Anthropic now adds a numerical watermark to outputs from models released after August 2, allowing Anthropic (though not necessarily end users) to determine whether Claude generated a given piece of text.
How did Fable 5.1 do on hands-on generation tests? #
Beyond benchmarks, informal side-by-side tests comparing Fable 5.1, GPT-5.6 Soul, and GLM 5.3 on creative and coding tasks showed Fable 5.1 performing well but not clearly winning. In a generated landscape scene test (rain clouds, a log cabin, a farm, a beach with boats and seagulls), Fable 5.1’s output was judged better than GLM 5.3’s but slightly behind GPT-5.6 Soul’s in visual polish.
Built like a system. Not vibe-coded.
Remy manages the project — every layer architected, not stitched together at the last second.
In a Rubik’s Cube simulation test, Fable 5.1 produced a notably feature-rich result, including labeled cube faces and granular customization sliders (individual color settings, sticker corner radius) not commonly seen in other models’ attempts. The simulation scrambled and solved correctly, matching the baseline capability of competing models while offering more configuration depth.
Frequently Asked Questions #
Is Fable 5.1 cheaper than Fable 5?
Anthropic claims yes, citing a 75% cut to cache read pricing and estimated savings of 25 to 45% depending on workload. However, independent analysis from Artificial Analysis found Fable 5.1 can cost more per completed task than Fable 5, because it uses about 1.7 times more output tokens to reach the same result.
Which model is the best value: Fable 5.1, GPT-5.6 Soul, or GLM 5.3?
For raw intelligence, Fable 5.1 currently scores highest on composite benchmarks. For cost efficiency, GPT-5.6 Soul offers close to frontier-level performance at a fraction of the price per task, making it the stronger choice for cost-sensitive, high-volume use.
What is the difference between Fable 5.1 and Mythos 5.1?
They are the same underlying model with different guardrail configurations. Mythos 5.1 consistently scores higher than Fable 5.1 on benchmarks like Terminal Bench 4 at equivalent reasoning effort, illustrating how added restrictions can reduce measured performance.
Does Fable 5.1 offer true zero data retention?
Not fully. Its new Enterprise Frontier Safeguards let customers control where data is stored, but Anthropic retains read access for misuse detection, which is a partial rather than complete fix for enterprise privacy concerns.
Why does Anthropic add watermarks to Fable 5.1’s output?
This is a compliance requirement under the EU AI Act’s code of practice on transparency for AI-generated content, applying to Anthropic models released after August 2. The watermark lets Anthropic determine the likelihood that Claude generated a given text.