{"slug": "when-an-ai-agent-says-done-and-it-is-not", "title": "When an AI agent says done and it is not", "summary": "A developer built 90 test traps in which a tool result only appeared to succeed, and found that Claude Haiku falsely reported tasks as complete in 56 of 272 runs while Claude Sonnet did so in 21 of 272 when given no extra instructions. The largest single cause was a write to a file or record that returned {\"ok\":true} but changed nothing, which neither model re-checked.", "body_md": "We built 90 traps where a tool result only looks like success. With no extra instructions, Claude Haiku claimed a false success in 56 of 272 runs.\n\nOften enough to matter, and nearly always in the same few places. We gave Claude Sonnet and Claude Haiku tasks whose tool results only looked like success. With no extra instructions, Sonnet reported work as finished that had never happened in 21 of 272 runs, and Haiku in 56 of 272. The single biggest source was a write to a file or a record that answered {\"ok\":true} and changed nothing, which neither model looked at again.\n\n**Read the full report on AISkills402:** [https://aiskills402.com/blog/agent-says-done-is-it](https://aiskills402.com/blog/agent-says-done-is-it)", "url": "https://wpnews.pro/news/when-an-ai-agent-says-done-and-it-is-not", "canonical_source": "https://dev.to/georgi_kalchev_85d01f7368/when-an-ai-agent-says-done-and-it-is-not-4k16", "published_at": "2026-10-08 07:06:31+00:00", "updated_at": "2026-10-08 07:17:12.002907+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "large-language-models", "ai-research"], "entities": ["Claude Haiku", "Claude Sonnet", "Anthropic", "AISkills402"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/when-an-ai-agent-says-done-and-it-is-not", "markdown": "https://wpnews.pro/news/when-an-ai-agent-says-done-and-it-is-not.md", "text": "https://wpnews.pro/news/when-an-ai-agent-says-done-and-it-is-not.txt", "jsonld": "https://wpnews.pro/news/when-an-ai-agent-says-done-and-it-is-not.jsonld"}}