{"slug": "why-does-opus-5-feel-worse-to-work-with", "title": "Why does Opus 5 feel worse to work with?", "summary": "Anthropic's Opus 5, despite being more capable than Opus 4.7 and Opus 4.8 and rivaling Fable in benchmarks, feels like a downgrade to work with because it makes assumptions and reinterprets plans without asking, according to an unnamed developer and colleagues. The author speculates that benchmark-driven training and the push for self-improving AI select for models that make bold assumptions rather than stopping to ask for clarification, which is problematic for real-world coding tasks where ambiguity is common.", "body_md": "[Why does Opus 5 feel worse to work with?](#why-does-opus-5-feel-worse-to-work-with)\n\nIn my opinion and that of the colleagues I've spoken with, working with Opus 5 feels like a downgrade compared to Opus 4.7, Opus 4.8, and Fable.\n\nI'm not claiming a step backwards in capabilities – it *is* a more capable model than Opus 4.7 and Opus 4.8 and even rivals\nFable in benchmarks, yet these other models feel better to work with. I believe this is because they:\n\n- stop and ask questions if my intent was unclear,\n- don't make assumptions without checking,\n- and don't reinterpret or update my plans without asking.\n\nBecause of this, they don't require the careful babysitting that Opus 5 does.\n\n[Baseless speculation](#baseless-speculation)\n\nI suspect this is the result of two compounding forces at Anthropic, and in current frontier labs in general.\n\nFirst, the desire to create a self-improving AI that is capable of recursively bootstrapping itself to AGI/ASI.\n\nSecond, the pressure to score highly on benchmarks. Although it's an open secret that many benchmark tasks are ill-defined, unfair, hackable, or otherwise broken, a good benchmark task is self-contained. It can be solved. It doesn't require hints, reading the task creator's mind, or outside information to pass.\n\nThat doesn't mean a good task can only have one correct answer, just that it should score *all* unambiguously correct\nanswers equally.\n\nSelecting for models that do well on benchmarks (and indeed training for them or on RLVR tasks in general) inherently selects for models that make bold, usually-correct assumptions in the face of ambiguity. It penalizes models with a tendency to stop and ask for clarification or direction.\n\nUnfortunately, that's exactly what most of us want from a coding agent.\n\nTry as you might, it's nearly impossible to get the entirety of the context, intentions, business implications, budget\nconstraints, and what-have-you written down and accessible to a coding agent. There will invariably\nbe ambiguity and choices to be made, and it is *nice* to know that an agent will stop and ask when needed.\n\nReal life just isn't a benchmark. There isn't a guaranteed right answer to every question, nor even a set of right answers, and with real-life consequences on the line, I do not want an agent taking its best guess!", "url": "https://wpnews.pro/news/why-does-opus-5-feel-worse-to-work-with", "canonical_source": "https://mun-logadan.github.io/why-does-opus-5-feel-worse/", "published_at": "2026-08-14 10:12:48+00:00", "updated_at": "2026-08-14 10:42:37.822571+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-products", "ai-research"], "entities": ["Anthropic", "Opus 5", "Opus 4.7", "Opus 4.8", "Fable"], "alternates": {"html": "https://wpnews.pro/news/why-does-opus-5-feel-worse-to-work-with", "markdown": "https://wpnews.pro/news/why-does-opus-5-feel-worse-to-work-with.md", "text": "https://wpnews.pro/news/why-does-opus-5-feel-worse-to-work-with.txt", "jsonld": "https://wpnews.pro/news/why-does-opus-5-feel-worse-to-work-with.jsonld"}}