Foresight AI Brings Gremlin Agents to Reliability Engineering Gremlin announced the general availability of Foresight AI, an agentic add-on to its reliability platform that analyzes services for potential failures, recommends changes and reruns tests to verify fixes. Foresight AI splits work across four agent roles — Analyst, Tester, Operator and Technical Program Manager — and bases recommendations on Gremlin's Failure Atlas, a proprietary record built from millions of fault-injection experiments over more than ten years across tens of thousands of distributed systems, with an LLM used only for searching, summarizing and explaining results. Gremlin CEO and co-founder Kolton Andrus said the product requires approval before any test runs or remediation is applied, keeping an engineer responsible for changes. Reliability management company Gremlin has announced the general availability of Foresight AI https://www.gremlin.com/blog/announcing-gremlin-foresight-ai , an agentic product that analyses services for potential failures, recommends changes and reruns tests to check that a fix works. It is available as an add-on to the Gremlin platform. The platform allows teams to run planned experiments, control the blast radius of failures, and use automated testing to ensure that fixes remain effective over time. By using reliability scores, organizations can quantify their progress and prioritise their investments in resilience. Foresight AI is intended to provide reliability checks without requiring every developer to become a distributed systems specialist. Gremlin positions the product as a response to faster AI-assisted software delivery, with Kolton Andrus, Gremlin's CEO and co-founder, arguing that more code is reaching production without close human review. "The best incident is the one that never happens." Kolton Andrus https://www.gremlin.com/blog/announcing-gremlin-foresight-ai The product divides work between four agent roles: an Analyst gathers service data and recommends tests, while a Tester schedules and runs them. An Operator interprets failed tests and proposes remediations. A Technical Program Manager tracks test coverage, reliability risks and commitments across teams. The latter can produce weekly reports and send updates through Slack. Foresight AI bases its recommendations on Gremlin's Failure Atlas. The company describes this as a proprietary record built from millions of fault-injection experiments conducted over more than ten years across tens of thousands of distributed systems. Gremlin says system data is not used to train a large language model. Instead, the agents use platform data and the Failure Atlas for decisions, with an LLM used for searching, summarising and explaining results. The main technical distinction is the closed validation loop. When a test uncovers a weakness, the service proposes a change and then runs the original test again. Gremlin says Foresight AI will neither run a test nor apply a remediation without approval. This keeps an engineer responsible for changes while automating much of the analysis and repetition around them. "An alert going quiet doesn't prove anything. A test that passes under load, in realistic conditions, does." Kolton Andrus https://www.gremlin.com/blog/announcing-gremlin-foresight-ai The service also creates reports and dashboards from natural-language requests. These can include charts, tables, tabs and links, and are intended to show changes in reliability scores and the risks that teams have addressed. This tackles a familiar problem for reliability teams: avoided incidents do not produce the same visible evidence as outages. In a separate article written for Gremlin https://www.gremlin.com/blog/why-ai-development-creates-a-reliability-blind-spot-for-humans-and-what-to-do-about-it , Intellyx analyst Jason English describes observability and AI SRE products as useful but largely reactive. He argues that they usually begin with telemetry from a problem that has already happened. Fault injection takes a different approach by creating controlled failure conditions before an incident and recording how a service responds. English also cautions that a reliability programme cannot simply be installed as a software package. Sustained results still need shared ownership and accountability across engineering and management. This qualification matters because Foresight AI can coordinate tests and commitments, but an organisation must still decide which risks are acceptable and whether proposed changes should proceed. Intellyx discloses that Gremlin is its customer. The release sits alongside other efforts to bring automated resilience checks into delivery pipelines, some of which have been covered recently by InfoQ. Google Cloud's chaos engineering guidance https://www.infoq.com/news/2025/11/google-chaos-engineering/ recommends realistic failure conditions, limited blast radii and continuous automation through CI/CD. Wix has also described using AI in CI/CD https://www.infoq.com/news/2025/07/wix-chaos-ai-cicd-pipelines/ for log interpretation and remediation suggestions, while keeping deployment and critical infrastructure decisions under human control. Those controls resemble the approval boundary in Foresight AI. They also differ from more autonomous models. Cloudflare's Agent Development Lifecycle proposal https://www.infoq.com/news/2026/09/cloudflare-adlc-agents/ , for example, envisages agents managing more of testing, deployment and maintenance. Gremlin currently takes a supervised approach, using agents to propose tests and changes while requiring people to authorise them. In a broader LinkedIn discussion about AI-based SRE automation https://www.linkedin.com/posts/sandipb site-reliability-engineers-and-other-operators-activity-7353486122083078145-3SDM , engineer Sandip Bhattacharya argued that organisations still need specialists to configure agents, assess their behaviour through metrics and adjust them as models and infrastructure change. This presents a useful test for Gremlin's supervised model: automation may reduce repetitive work, but it also creates a new system that reliability engineers must evaluate. Mike Dauber, a general partner at Amplify Partners, compared Foresight AI to "a trainer who does the reps for you" https://www.prnewswire.com/news-releases/gremlin-launches-foresight-ai-to-proactively-fix-reliability-risks-302898275.html . His comment captures the product's emphasis on repeated testing, although the evidence presented so far comes from Gremlin and its beta programme. The company has not published comparative accuracy data, false-positive rates or detailed beta results. Gremlin was founded by engineers with experience at Amazon and Netflix. https://www.gremlin.com/about Andrus previously worked on reliability at both companies, while Netflix's Chaos Monkey helped popularise automated failure injection. Foresight AI extends that approach from running experiments towards selecting tests, interpreting their results, proposing fixes and checking those fixes over time. This release follows a successful beta period where the technology demonstrated an ability to identify and address reliability risks before they cause incidents, and Gremlin Foresight AI is available for general use starting today.