Member-only story
Anthropic Secretly Nerfed Claude for Six Weeks. Then They Admitted It.
Reasoning effort dropped from HIGH to MEDIUM on March 4. A caching bug deleted reasoning history on March 26. Benchmark accuracy fell 18 points. Anthropic only disclosed it at the Opus 4.8 launch on May 28. Here is the paper trail.
On March 4, 2026, Claude got quietly dumber.
Not catastrophically. Not obviously. But for developers running complex reasoning tasks — multi-step problem solving, code review, mathematical proofs, long document analysis — something measurably changed.
The model started giving shorter chains of thought. It started skipping verification steps it used to do automatically. It started being wrong more often on hard problems.
Nobody at Anthropic said anything.
The developers who noticed started posting on the Anthropic developer forums.
They ran their own evals. They compared outputs. They opened support tickets.
Anthropic’s response, for six weeks: nothing official.
Then, on May 28, 2026 — at the launch event for Opus 4.8 — Anthropic’s engineering team quietly acknowledged what had happened.
In a technical disclosure buried in the release notes, they described a two-part…