Long-horizon browser-based computer-use benchmark built from deterministic enterprise-web environments. Its five headline slices score pass@1 against certified optimal solutions, while a sixth number-partitioning slice reports a separate reward ladder for deliberately intractable tasks. Category: Agentic. Imported rows: 8. Top imported result: Claude Fable 5, rank 1, 74.
Debate-Driven Development: Why AI Agents That Argue Over Your Code Catch 30% More Bugs