Human beings aren’t perfect, so it should be no surprise that any AI system designed and coded by humans can go wrong.
Last July, Jason Lemkin, a tech entrepreneur and founder of the SaaS community SaaStr, documented his experiment with Replit's AI-assisted “vibe coding” tool, which managed to delete a production database during a standard code freeze. An AI developer agent bypassed that code freeze, generated fake user records to hide its mistake, and wiped a live production database after it incorrectly triggered a command.
As we learned from the OpenAI-Hugging Face incident, there are still a lot of things that enterprise tech does not understand about AI agent behavior. The incident highlighted the need for a precise system rollback tool to undo an AI agent’s specific actions – in Replit's case, developers were forced to manually restore old backups, and that can cause a massive amount of operational downtime.
According to reports at the time, Lemkin was able to manually roll back the damage caused by the AI agent within a couple of hours. Replit responded immediately too, with Replit CEO Amjad Masad publicly apologising on X and calling the incident "unacceptable."
In a separate incident later that year, Amazon's Kiro AI autonomous coding tool accidentally bypassed approval steps and wiped a production environment, which it also decided to recreate, which caused a 13-hour outage for the Cost Explorer service in China. Amazon attributed the incident to user error and misconfigured access controls, saying that the engineer had granted the agent overly broad permissions.
Both events highlighted the importance of traceability when working with AI agents, which is easier said than done.
All about trust
Get the full story: Subscribe for free #
Join peers managing over $100 billion in annual IT spend and subscribe to unlock full access to The Stack’s analysis and events.
[Subscribe now](https://www.thestack.technology/membership/)
Already a member? [Sign in](https://www.thestack.technology/signin/)