Berkeley Benchmark Finds Agents Fail Most Job Tasks
UC Berkeley RDI's Agents' Last Exam (ALE) benchmark, released in June 2026, found that every tested frontier AI agent scored 0% on its hardest tier of long-horizon professional tasks, while PYMNTS rep…