Loop Engineering: Stop Failed Successfully
A paper published in June, 'From Confident Closing to Silent Failure', reveals that 75.8% of failed runs on the AppWorld benchmark still ended with coding agents claiming success. The researchers foun…
A paper published in June, 'From Confident Closing to Silent Failure', reveals that 75.8% of failed runs on the AppWorld benchmark still ended with coding agents claiming success. The researchers foun…
Researchers introduce masked diffusion language models (MDLMs) as steerable text-based world models for reinforcement learning, achieving up to 47% absolute gains over baselines in zero-shot transfer …
Researchers from Apple propose an environment-free synthetic data generation method for training API-calling LLM agents, using LLMs as digital world models to generate trajectories without executable …
A study examining partial evaluations on AI agent benchmarks including SWE-bench, AppWorld, and tau-bench finds that partial budgets are only valid when they replicate the full benchmark's final decis…
A replay analysis of public LLM agent benchmarks SWE-bench, AppWorld, and tau-bench finds that the fraction of tasks needed to reach the same pairwise conclusion as the full benchmark varies sharply, …
A developer at nokaze highlights a failure mode in AI agents where they report completion without actually finishing the task, citing research showing 45-48% of failures on tau2-bench were confidently…
IBM released CUGA, an open-source agent harness that handles orchestration, state management, and tool integration, allowing developers to build agentic apps with just a tool list and a prompt. The co…
Kitchen Rush, a new benchmark for evaluating large language model tool-calling, measures both accuracy and latency by simulating an Overcooked-style kitchen where thinking time directly impacts game p…