{"slug": "google-s-open-mantis-kit-helps-coding-agents-find-and-patch-security-bugs", "title": "Google's open Mantis kit helps coding agents find and patch security bugs", "summary": "Google released Mantis, an open-source kit of skills and an ADK reference harness that lets coding agents autonomously find, triage, reproduce, and patch security vulnerabilities in codebases. Mantis installs via a reference install script and is driven by mantis-configure and mantis-launch tools, with commands such as ./run.sh path/to/code, a --focus flag for targeting bug classes like IDOR, and --objective for research graph synthesis. Google warns that the AI models are non-deterministic and can hallucinate findings or generate incorrect patches, so all findings must be manually verified by a security expert and the suite should only run in isolated, restricted environments.", "body_md": "[!CAUTION] **USE AT YOUR OWN RISK. BE EXTREMELY CAREFUL.** This suite is\ndesigned to generate and execute autonomously generated code that may be\nunstable or perform unexpected actions. **USE THIS ONLY IN ISOLATED,\nRESTRICTED ENVIRONMENTS.** Never run this suite on a machine with access to\nproduction systems, sensitive data, or internal networks.\n\n[!IMPORTANT] **RESPONSIBLE USE** AI models are non-deterministic and can\nhallucinate findings or generate incorrect patches. **All findings must be\nmanually verified by a security expert before being reported.** Do not\nmass-file unverified, AI-generated reports to open-source maintainers. A\nfailure to automatically reproduce a vulnerability does not definitively mean\nit is a false positive, nor does a successful reproducer guarantee the bug is\nexploitable in all contexts. Use Mantis responsibly.\n\nMantis is a set of skills along with an ADK reference harness for building secure software in the new AI era of software development.\n\nFirst, install python3-venv such as with `sudo apt install python3-venv`, then\nrun the install script. Mantis comes with automated configuration and launcher\ntools (`mantis-configure` and `mantis-launch`):\n\n```\ncd reference && ./install.sh\n\n# 0. Authenticate Google Cloud Application Default Credentials (ADC) if using Vertex AI\ngcloud auth application-default login\n\n# 1. Fast Configuration & Capability Auto-Detection (or --interactive wizard)\npython3 scripts/configure.py --auto\n\n# 2. Fast Preflight Validation (~1s) & Live Reachability Probe\npython3 scripts/configure.py --test --probe\n\n# 3. Launch Vulnerability Review Campaign (file or repository)\n./run.sh path/to/code            # a file or a directory\n\n# 4. (Optional) Tell the planner what to hunt for, in plain language (standard pipeline)\n./run.sh path/to/code --focus \"look for IDOR\"\n\n# 5. (Optional) Remove spend limits for long unattended runs\n./run.sh path/to/code --no-budget --parallel 32\n\n# 6. (Optional) Research Graph Synthesis: designs a new agent-graph topology for your objective\n./run.sh path/to/code --objective \"Audit for Server-Side Request Forgery and SSRF in webhook handlers\"\n```\n\nMantis is roughly designed to:\n\n- Review history of a codebase to look for historical vulnerabilities we do not wish to repeat\n- Build a semantic index (and/or summaries) of the codebase for efficient code navigation\n- Automatically generate a threat model\n- Build up a set of hypotheses for individual agents to research (or simply \"scan every file\")\n- Execute on those research plans\n- Deduplicate existing findings\n- Triage/critique findings to combat hallucination and statically verify production viability\n- Reproduce the vulnerability to varying degrees, depending on available environments, anywhere from a static guess at what an exploit might look up up to a unit test or even spinning up a mock server to attempt to exploit\n- Look across known vulnerabilities to attempt to build more impactful chains\n- Patch discovered vulnerabilities, using an adversarial loop to verify the vulnerability is really fixed\n- Calibrate all findings based on an established rubric to combat LLM inflation of severity (surfacing the most critical risks to humans instead of spewing thousands of \"criticals\")\n- Reflect on each iteration of the loop to look for things we've learned in the now completed round of research\n- Based on all of the collected learnings, threat model, and knowledge base, use\nthe `/mantis-advise` skill to develop code more securely and ensure that\nduring development you do not repeat prior mistakes or trigger edge cases in\ncode that were previously protected by some guard that was removed\n\nAdditionally, the ADK reference harness shows some of the neat ways in which we\ncan build agentic workflows. Specifically, `reference/workflow.json` shows how a\ndeep review is done, and research graph synthesis\n(`./run.sh target --objective \"...\"`) allows generating custom graph topologies\non the fly. This is a very powerful construct because you can use this to build\nany kind of agentic deep dive security review you might imagine.\n\nMantis is intended to be a starting point rather than a rigid set of instructions. You should adapt, tune, and extend this harness to fit your organization's specific software or hardware stack. We provide an ADK-based reference harness which works out of the box, but any competent coding agent should be able to convert this to use your framework of choice.\n\nThe Mantis skills can be adapted to\n[specialized domains](https://github.com/google/mantis/blob/main/README_AGENTS.md#adaptability--specialized-domains) (such\nas Hardware/RTL, Infrastructure as Code, ML pipelines, or compiled firmware).\n\nWe strongly recommend using AI to iterate on these skills and using your internal documentation, coding standards, and build systems to augment the threat models you use for scanning. We also strongly recommend adapting risk calibration to your environment and risk tolerance.\n\nAbove all, while orchestrated vulnerability discovery is incredibly powerful and useful, it is even more important to use this in a suitably isolated environment to prevent impacting production systems. The ADK reference harness provides some example sandboxes for testing, but if you use the most advanced frontier models you should go beyond this and additionally set up an additional sandboxing layer that itself contains very strong monitoring to look for escape attempts.\n\nThis is not an officially supported Google product. This project is not eligible\nfor the\n[Google Open Source Software Vulnerability Rewards Program](https://bughunters.google.com/open-source-security).\n\nThis project is intended for demonstration purposes only. It is not intended for use in a production environment.", "url": "https://wpnews.pro/news/google-s-open-mantis-kit-helps-coding-agents-find-and-patch-security-bugs", "canonical_source": "https://github.com/google/mantis", "published_at": "2026-10-02 14:31:39+00:00", "updated_at": "2026-10-02 14:38:55.365327+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "developer-tools", "ai-tools", "artificial-intelligence"], "entities": ["Google", "Mantis", "ADK", "Vertex AI", "Google Cloud"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/google-s-open-mantis-kit-helps-coding-agents-find-and-patch-security-bugs", "markdown": "https://wpnews.pro/news/google-s-open-mantis-kit-helps-coding-agents-find-and-patch-security-bugs.md", "text": "https://wpnews.pro/news/google-s-open-mantis-kit-helps-coding-agents-find-and-patch-security-bugs.txt", "jsonld": "https://wpnews.pro/news/google-s-open-mantis-kit-helps-coding-agents-find-and-patch-security-bugs.jsonld"}}