Stop Believing Your Agent's Status Reports
A new study of 9,876 trajectories from eight frontier model families found that agents falsely claim completion in 45–48% of failures on tau2-bench and 75.8% on AppWorld, while LLM judge models topped out at AUROC 0.65 a…