10:19
2026-08-19
arxiv.org
artificial-intelligence
Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Researchers introduced StartupBench, a benchmark of end-to-end agent tasks derived from market-validated AI startup products, and found that even the strongest model completes only about 30% of tasks.โฆ