{"slug": "claude-did-best-on-a-new-benchmark-for-agents-that-build-agents-it-still-passed", "title": "Claude did best on a new benchmark for ‘agents that build agents’. It still passed fewer than a quarter of the tests.", "summary": "Anthropic's Claude topped a new benchmark for AI agents that build other agents, yet it still passed fewer than a quarter of the tests, according to The New Stack. The benchmark evaluates models' ability to create functional agents, highlighting the current limitations of even the best-performing AI systems.", "body_md": "AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems\n\nThe post [Claude did best on a new benchmark for ‘agents that build agents’. It still passed fewer than a quarter of the tests.](https://thenewstack.io/claude-build-agents-benchmark/) appeared first on [The New Stack](https://thenewstack.io).", "url": "https://wpnews.pro/news/claude-did-best-on-a-new-benchmark-for-agents-that-build-agents-it-still-passed", "canonical_source": "https://thenewstack.io/claude-build-agents-benchmark/", "published_at": "2026-09-09 20:14:09+00:00", "updated_at": "2026-09-09 23:14:53.840509+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research"], "entities": ["Anthropic", "Claude", "The New Stack"], "alternates": {"html": "https://wpnews.pro/news/claude-did-best-on-a-new-benchmark-for-agents-that-build-agents-it-still-passed", "markdown": "https://wpnews.pro/news/claude-did-best-on-a-new-benchmark-for-agents-that-build-agents-it-still-passed.md", "text": "https://wpnews.pro/news/claude-did-best-on-a-new-benchmark-for-agents-that-build-agents-it-still-passed.txt", "jsonld": "https://wpnews.pro/news/claude-did-best-on-a-new-benchmark-for-agents-that-build-agents-it-still-passed.jsonld"}}