cd /news/artificial-intelligence/why-does-opus-5-feel-worse-to-work-w… · home topics artificial-intelligence article
[ARTICLE · art-96648] src=mun-logadan.github.io ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

Why does Opus 5 feel worse to work with?

Anthropic's Opus 5, despite being more capable than Opus 4.7 and Opus 4.8 and rivaling Fable in benchmarks, feels like a downgrade to work with because it makes assumptions and reinterprets plans without asking, according to an unnamed developer and colleagues. The author speculates that benchmark-driven training and the push for self-improving AI select for models that make bold assumptions rather than stopping to ask for clarification, which is problematic for real-world coding tasks where ambiguity is common.

read2 min views2 publishedAug 14, 2026

Why does Opus 5 feel worse to work with? In my opinion and that of the colleagues I've spoken with, working with Opus 5 feels like a downgrade compared to Opus 4.7, Opus 4.8, and Fable.

I'm not claiming a step backwards in capabilities – it is a more capable model than Opus 4.7 and Opus 4.8 and even rivals Fable in benchmarks, yet these other models feel better to work with. I believe this is because they:

  • stop and ask questions if my intent was unclear,
  • don't make assumptions without checking,
  • and don't reinterpret or update my plans without asking.

Because of this, they don't require the careful babysitting that Opus 5 does.

Baseless speculation I suspect this is the result of two compounding forces at Anthropic, and in current frontier labs in general.

First, the desire to create a self-improving AI that is capable of recursively bootstrapping itself to AGI/ASI.

Second, the pressure to score highly on benchmarks. Although it's an open secret that many benchmark tasks are ill-defined, unfair, hackable, or otherwise broken, a good benchmark task is self-contained. It can be solved. It doesn't require hints, reading the task creator's mind, or outside information to pass.

That doesn't mean a good task can only have one correct answer, just that it should score all unambiguously correct answers equally.

Selecting for models that do well on benchmarks (and indeed training for them or on RLVR tasks in general) inherently selects for models that make bold, usually-correct assumptions in the face of ambiguity. It penalizes models with a tendency to stop and ask for clarification or direction.

Unfortunately, that's exactly what most of us want from a coding agent.

Try as you might, it's nearly impossible to get the entirety of the context, intentions, business implications, budget constraints, and what-have-you written down and accessible to a coding agent. There will invariably be ambiguity and choices to be made, and it is nice to know that an agent will stop and ask when needed.

Real life just isn't a benchmark. There isn't a guaranteed right answer to every question, nor even a set of right answers, and with real-life consequences on the line, I do not want an agent taking its best guess!

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/why-does-opus-5-feel…] indexed:0 read:2min 2026-08-14 ·