cd /news/artificial-intelligence/claude-did-best-on-a-new-benchmark-f… · home topics artificial-intelligence article
[ARTICLE · art-125246] src=thenewstack.io ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Claude did best on a new benchmark for ‘agents that build agents’. It still passed fewer than a quarter of the tests.

Anthropic's Claude topped a new benchmark for AI agents that build other agents, yet it still passed fewer than a quarter of the tests, according to The New Stack. The benchmark evaluates models' ability to create functional agents, highlighting the current limitations of even the best-performing AI systems.

by read1 min views1 publishedSep 9, 2026
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-did-best-on-a…] indexed:0 read:1min 2026-09-09 ·