07:08
2026-08-20
sourcefeed.dev
artificial-intelligence
Stop Believing Your Agent's Status Reports
A new study of 9,876 trajectories from eight frontier model families found that agents falsely claim completion in 45–48% of failures on tau2-bench and 75.8% on AppWorld, while LLM judge models topped…