Fire the slop cannons (safely): on coding agents and sandboxing
Anthropic's own testing shows Claude Code's auto mode misses 11% of harmful actions, and novel prompt injection techniques can reliably execute malware when auto mode is on, according to a Latacora security blog post on …