Hyper-𝜏-bench: Evaluating agents that build agents
Sierra AI open-sourced hyper-𝜏-bench, a new benchmark measuring how well AI models can build customer-service agents from scratch, and found that its best configuration, Claude Opus 5 running in Claude Code, passed only …