# Claude did best on a new benchmark for ‘agents that build agents’. It still passed fewer than a quarter of the tests.

> Source: <https://thenewstack.io/claude-build-agents-benchmark/>
> Published: 2026-09-09 20:14:09+00:00

AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems

The post [Claude did best on a new benchmark for ‘agents that build agents’. It still passed fewer than a quarter of the tests.](https://thenewstack.io/claude-build-agents-benchmark/) appeared first on [The New Stack](https://thenewstack.io).
