cd /news/artificial-intelligence/anthropic-says-claude-is-showing-ear… · home topics artificial-intelligence article
[ARTICLE · art-114732] src=startupfortune.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Anthropic Says Claude Is Showing Early Signs of Self-Improvement

Anthropic reports that over 80% of code merged into its codebase as of May 2026 was authored by Claude, up from low single digits before Claude Code's preview in February 2025. The company says Claude's success rate on open-ended coding tasks reached 76% in May 2026, a 50-percentage-point increase in six months, and that Claude Mythos Preview achieved a 52x speedup on training-code optimization by April 2026. Anthropic co-founder Jack Clark and Anthropic Institute lead Marina Favaro co-authored the analysis, which also notes Claude-powered agents recovered 97% of the gap in a safety research test over 800 hours using about $18,000 in compute.

read6 min views3 publishedAug 28, 2026
Anthropic Says Claude Is Showing Early Signs of Self-Improvement
Image: Startupfortune (auto-discovered)

Anthropic has put a number on a problem most AI labs still talk about carefully: Claude is already helping build Claude, and the hard part now is not code. It is judgment.

Anthropic's own evidence is blunt. In its Institute piece, "When AI builds itself," co-authored by Anthropic co-founder Jack Clark and Anthropic Institute lead Marina Favaro, the company says more than 80% of the code merged into its codebase as of May 2026 was authored by Claude. Before Claude Code entered research preview in February 2025, that figure was in the low single digits.

That is not a small productivity story. It is the start of the industry's most uncomfortable loop. The people building frontier AI are now using frontier AI to build more of the next system, and Anthropic is saying the quiet part in public.

The company also says its engineers were merging 8 times as much code per day in the second quarter of 2026 as they were in 2024, while warning that lines of code overstate true productivity. Good. That caveat belongs there. More code is not the same thing as better software, and anyone who has reviewed a sprawling automated pull request knows the difference.

Still, the direction is clear. Anthropic says Claude's success rate on its most open-ended coding tasks reached 76% in May 2026, up 50 percentage points in six months. One internal example involved a routine upgrade that was crashing tens of thousands of training jobs. An engineer gave Claude cluster access and little more than text context, and Claude isolated the obscure debugging flag behind the crash in about two hours. Anthropic says the same work would normally take two to three days.

Anthropic’s Mythos puts AI engineering economics under pressure Claude Mythos Preview is still restricted, but its reported 52x speedup on training-code optimization points to a new phase in AI-assisted engineering. For founders, the real question is how quickly faster technical iteration turns into lower costs, tighter workflows and new access divides. - how Mythos AI training speedup affects startup economics - what happens when AI engineering costs drop dramatically

That is the part you should watch.

The Code Is No Longer The Boundary #

Anthropic's strongest evidence is not the 80% code figure by itself. It is the move from writing code to running experiments. The company describes a standing internal test where Claude is given code that trains a small model and is asked to make it run faster without breaking correctness checks. Claude Opus 4 averaged about a 3x speedup in May 2025. By April 2026, Claude Mythos Preview was reaching about 52x.

That kind of result changes the job. You are no longer asking whether a model can autocomplete a function or clean up boilerplate. You are asking whether it can grind through the experimental loop faster than a human researcher can, while the human decides what question was worth asking in the first place.

Anthropic's open-ended safety research test gets closer to the live issue. The company says Claude-powered agents were given a problem about whether a weaker model can reliably supervise a stronger one, then left to propose hypotheses, run tests, share findings and iterate. Two human researchers recovered about 23% of the gap between the weak model's floor and the stronger model's ceiling over roughly a week. The agents recovered 97% over more than 800 cumulative hours, using about $18,000 in compute.

There are limits. Anthropic says the result did not transfer cleanly to production-scale models, and humans still chose the problem and scoring rubric. That is not a footnote. That is the whole fight.

Claude is becoming very good at execution. It has not yet taken ownership of taste, direction and judgment. Those are different muscles.

Rivals Are Pushing The Same Line #

Anthropic is not alone here. OpenAI said in February that early versions of GPT-5.3-Codex helped debug the training run for that release, manage deployment work and diagnose evaluations. In July, OpenAI said GPT-5.6 had helped make itself more efficient to run before the company cut prices for parts of the GPT-5.6 family.

Claude Mythos is turning AI benchmarks into a founder question Claude Mythos Preview is being shared as a 17-hour AI task-horizon story, but METR's own warning makes the number less precise than it sounds. The real issue for founders is how quickly longer autonomous work becomes verifiable, affordable and safe inside actual startup workflows. - how to verify AI agent work for startups - what happens to contract coding with AI agents

Google DeepMind has its own version of the story. Its AlphaEvolve system has been used to optimize Google's infrastructure, including next-generation TPU design, Spanner storage behavior and Gemini training kernels. DeepMind said AlphaEvolve found a 23% speedup for a Gemini architecture kernel, which translated into a 1% reduction in Gemini training time. A single percent sounds dull until you remember what frontier training runs cost.

So no, this is not only Anthropic being dramatic. The biggest labs are all discovering that AI can improve the machinery around AI. The difference is that Anthropic is framing it directly as progress toward recursive self-improvement, while also spelling out the risk of losing meaningful human control if the loop closes too far.

The Skeptics Have A Point #

The cleanest challenge came this month. MIT Technology Review reported on August 18 that a Princeton-led team including Peter Kirgis and Sayash Kapoor tested Claude Opus 4.8 on two unpublished NeurIPS 2026 research questions using a method called shadow evaluation. The agents had six days, $3,000 in Anthropic API credits, GPU access and the open web. The original authors rejected both AI-produced papers.

Kapoor told MIT Technology Review the work was "nowhere close to the mark" for a top AI conference. Jack Clark, writing in his Import AI newsletter, called the findings a "bearish signal" for short recursive self-improvement timelines and pointed to models still struggling with creative research taste.

That skepticism is not anti-AI. It is useful. Models can do well when the answer can be checked automatically, as in code speedups, benchmark optimization and many engineering tasks. Open-ended research is messier. You have to know when a result is boring, when a method is broken, when an apparent failure is actually the interesting thing, and when to stop digging.

Here's the thing: Anthropic and its critics are not really disagreeing on the facts. Claude is already writing most of Anthropic's code. It is already improving training code and running research loops under human direction. It is also still weak at choosing the right big question. If you run a company using these tools, that distinction matters more than the phrase recursive self-improvement itself.

The next Claude may not be built by Claude alone. But it will almost certainly be built with Claude doing more of the work than it did last year, and with humans pushed further toward the smaller, harder job of deciding what should be built at all.

Also read: Fed Chair Kevin Warsh's Inflation Warning Rattles Tech and AI StocksNvidia's $12.9 Billion Bid for Hugging Face Signals a New AI Land GrabSouth Korea Picks SKT, KT and Kakao to Build Its Free National AI Chatbot

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-says-claud…] indexed:0 read:6min 2026-08-28 ·