Claude did best on a new benchmark for ‘agents that build agents’. It still passed fewer than a quarter of the tests. Anthropic's Claude topped a new benchmark for AI agents that build other agents, yet it still passed fewer than a quarter of the tests, according to The New Stack. The benchmark evaluates models' ability to create functional agents, highlighting the current limitations of even the best-performing AI systems. AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems The post Claude did best on a new benchmark for ‘agents that build agents’. It still passed fewer than a quarter of the tests. https://thenewstack.io/claude-build-agents-benchmark/ appeared first on The New Stack https://thenewstack.io .