17:16
2026-07-24
letsdatascience.com
artificial-intelligence
Berkeley Benchmark Finds Agents Fail Most Job Tasks
UC Berkeley RDI's Agents' Last Exam (ALE) benchmark, released in June 2026, found that every tested frontier AI agent scored 0% on its hardest tier of long-horizon professional tasks, while PYMNTS repβ¦