You can restart an AI agent. Can you undo what it did? Replit's AI-assisted 'vibe coding' tool deleted a production database during a code freeze in July, forcing a manual rollback that took hours, while Amazon's Kiro AI tool caused a 13-hour outage for its Cost Explorer service in China by bypassing approval steps and wiping a production environment. These incidents underscore the need for precise rollback tools and traceability in AI agent operations, as highlighted by tech entrepreneur Jason Lemkin and Replit CEO Amjad Masad, who apologized publicly. Agents /tag/agents/ Human beings aren’t perfect, so it should be no surprise that any AI system designed and coded by humans can go wrong. Last July, Jason Lemkin, a tech entrepreneur and founder of the SaaS community SaaStr, documented his experiment https://www.thestack.technology/vibe-coding-ceo-deletes-production-database/ with Replit's AI-assisted “vibe coding” tool, which managed to delete a production database during a standard code freeze. An AI developer agent bypassed that code freeze, generated fake user records to hide its mistake, and wiped a live production database after it incorrectly triggered a command. As we learned from the OpenAI-Hugging Face incident, there are still a lot of things that enterprise tech does not understand about AI agent behavior. The incident highlighted the need for a precise system rollback tool to undo an AI agent’s specific actions – in Replit's case, developers were forced to manually restore old backups, and that can cause a massive amount of operational downtime. According to reports at the time, Lemkin was able to manually roll back the damage caused by the AI agent within a couple of hours. Replit responded immediately too, with Replit CEO Amjad Masad publicly apologising on X https://x.com/amasad/status/1946986468586721478?ref=thestack.technology and calling the incident "unacceptable." In a separate incident later that year, Amazon's Kiro AI autonomous coding tool accidentally bypassed approval steps and wiped a production environment, which it also decided to recreate, which caused a 13-hour outage for the Cost Explorer service in China. Amazon attributed the incident to user error and misconfigured access controls, saying that the engineer had granted the agent overly broad permissions. Both events highlighted the importance of traceability when working with AI agents, which is easier said than done. All about trust Get the full story: Subscribe for free Join peers managing over $100 billion in annual IT spend and subscribe to unlock full access to The Stack’s analysis and events. Subscribe now https://www.thestack.technology/membership/ Already a member? Sign in https://www.thestack.technology/signin/