cd/entity/Agents' Last Exam· home entities Agents' Last Exam
grep -l @agents' last exam /news/*.json | wc -l → 10

Agents' Last Exam

mentions 10 type Person feed RSS

// recent coverage 10 mentions

17:16
2026-07-24
letsdatascience.com
artificial-intelligence

Berkeley Benchmark Finds Agents Fail Most Job Tasks

UC Berkeley RDI's Agents' Last Exam (ALE) benchmark, released in June 2026, found that every tested frontier AI agent scored 0% on its hardest tier of long-horizon professional tasks, while PYMNTS rep…

23:08
2026-07-15
byteiota.com
large-language-models

GPT-5.6 Sol, Terra, Luna: Which Model for Your Stack

OpenAI shipped GPT-5.6 on July 9 as three distinct models — Sol, Terra, and Luna — not a single upgrade, with the gpt-5.6 alias routing to Sol at $5 per million input tokens. The release introduces Pr…

19:46
2026-07-09
simonwillison.net
artificial-intelligence

The new GPT-5.6 family: Luna, Terra, Sol

OpenAI released the GPT-5.6 family of models in three sizes—Luna, Terra, and Sol—claiming superior long-running agentic performance over Anthropic's Claude Fable 5 on the Agents' Last Exam benchmark, …

14:39
2026-06-15
letsdatascience.com
ai-agents

Agents' Last Exam launches economically focused agent benchmark

Berkeley RDI and over 300 industry experts launched Agents' Last Exam (ALE), an open benchmark that evaluates AI agents on economically valuable professional tasks rather than abstract proxies. The be…

15:00
2026-06-14
nlp.elvissaravia.com
artificial-intelligence

🥇Top AI Papers of the Week

Researchers at MiniMax introduced MiniMax Sparse Attention (MSA), a method that reduces attention compute by 28.4x at 1M context while matching performance, enabling cheaper long-context deployment. B…

00:00
2026-06-10
epics.tech
artificial-intelligence

Fable 5 and the Async-Agent Era

Anthropic launched Fable 5 / Mythos 5, a model tier for long-horizon asynchronous autonomy, on June 10, 2026, marking the arrival of the async-agent era. The release reframes AI competition around aut…

04:00
2026-06-06
arxiv.org
artificial-intelligence

Agents' Last Exam

Researchers introduced Agents' Last Exam (ALE), a new benchmark designed to evaluate AI agents on long-horizon, economically valuable, real-world tasks with verifiable outcomes. Developed with over 25…

// co-occurs with top 8 entities
// topics top 6 topics