Gremlin now uses AI to break distributed systems faster Gremlin launched Foresight AI, an add-on to its chaos engineering services that autonomously injects failures into distributed systems, identifies root causes, and generates fixes, according to founder Kolton Andrus. The tool is grounded in Gremlin's Failure Atlas, a repository of millions of chaos engineering experiments run over the past decade on "tens of thousands of systems," which the company says keeps the underlying LLM on point and minimizes hallucinations. Gremlin's existing guardrails limit the blast radius using enterprise access permissions, and Foresight AI can either apply its proposed code or configuration change or prepare a report for an SRE, with humans kept in the loop at critical junctures. Gremlin now uses AI to break distributed systems faster Source: The Register https://www.theregister.com If there's one thing AI can do well, it's failing more quickly than the average SRE What better way to find out what your system can take than by breaking it? Popularized by Netflix more than a decade ago, chaos testing is the practice of deliberately introducing controlled failures into working systems to test their resilience. Now, a startup that built its business around this practice is using AI to automate more of the testing and troubleshooting process. Gremlin’s Foresight AI, a new add-on to its current “chaos engineering” services, autonomously breaks software infrastructure to uncover hidden bugs before they pop up in production. Chaos engineering can be deliciously dangerous for an engineer: Turn off a random server, load balancer or even an entire region of a cloud deployment, then observe how well the rest of the system tries to keep serving customers. If it does, job well done. If it's borked, more reliability work is afoot. Kolton Andrus, one of the Netflix engineers who did chaos engineering at the streaming service, founded Gremlin to bring fault injection to enterprises, along with business-friendly accoutrements such as scoring, executive reporting, and organizational accountability tools. Adding AI into the mix speeds the process considerably, Andrus told The Register, by automating many of the preparatory and post-testing tasks that used to be done by hand. Failure-as-a-Service AIOps services are nothing new for the site reliability engineer SRE , but most of these tools diagnose issues only after the system has already crashed. "A lot of AI solutions are, 'Hey, we took a guess. Here you go. Good luck,'” Andrus said. Instead, Gremlin purposefully tips over the system and then offers expert advice on how to prevent that from happening again. To use Gremlin, the user installs agents on their system that can inject latency, kill pods, or max out memory. A Gremlin cloud service then observes the subsequent telemetry alerts for telltale signs of systemic weakness. Gremlin’s secret sauce is its Failure Atlas, a repository of millions of chaos engineering experiments that the company conducted over the past decade on “tens of thousands of systems,” Andrus said. Gremlin developed a harness that yokes the LLM https://www.machinebrief.com/glossary/llm to the Failure Atlas, allowing it to generate more technically grounded answers about what went wrong, rather than simply drawing on its own collection of random information found on the Internet, the company asserts. Gremlin uses a variety of closed and open LLMs for the job, depending on which one is working best at the time. The Atlas keeps the LLM on point, minimizing hallucinations, Andrus said. Not creating problems, just pointing them out Gremlin’s existing guardrails https://www.machinebrief.com/glossary/guardrails are built around the enterprise’s own access permissions and other security precautions to keep the “blast radius” company’s words to a minimum if the blast isn’t contained at all, well, that’s an access control issue right there . Once a system failure is induced, Foresight AI determines the root cause, generates code or the configuration change it thinks will fix the issue, and then can either apply the solution, or prepare a report for an SRE to read. It’s basically an agentic loop of test-and-replace – keeping humans in the loop at critical junctures – that runs until the failure is no longer induced. Gremlin’s current user base tends to be larger compute https://www.machinebrief.com/glossary/compute -heavy enterprises, many specializing in financial services, retail and enterprise SaaS, Andrus said. They use Gremlin to test disaster recovery plans, ensure correct operation under less-than-ideal conditions, and make sure Kubernetes scales correctly. Such complex distributed systems are particularly difficult to diagnose when they go tits up. Difficult to catch with routine integration and unit testing, some of these bugs can be flushed out of hiding by injecting faults such as network latency, memory exhaustion, or other conditions that expose weaknesses in distributed systems. The corporate rush to use AI to both write applications and deploy infrastructure brings its own set of complications, explained Andrus in an interview with The Register. AI works at such a speed that it is impossible for human reviewers to keep up, and AI can make dumb mistakes anywhere. Any service that promises to break systems for the greater good will require sign-off at the highest corporate levels, said noted Intellyx analyst Jason English, in a note about the technology. For it to work effectively, “we need to change our mindset about risk,” he wrote.® Get AI news in your inbox Daily digest of what matters in AI.