{"slug": "more-than-half-of-ai-generated-patches-are-broken", "title": "More than half of AI-generated patches are broken", "summary": "New research from 1Password found that OpenAI's ChatGPT 5.5 and Anthropic's Claude Opus 4.8 successfully patched only 47% of six high-impact, high-complexity CVEs, often leaving vulnerabilities partially addressed or introducing new bugs. A separate Veracode report testing 100 models found an average security pass rate of 56% for AI-generated code, with 44% of tests introducing an OWASP Top 10 vulnerability. The findings suggest that autonomous AI patching is not yet reliable without human oversight.", "body_md": "# More than half of AI-generated patches are broken\n\nAs AI-generated code continues to be injected into all corners of the internet, concerns have risen about an expanding attack surface for malicious hackers to exploit.\n\nSome have argued that the enhanced cybersecurity capabilities of large language models could serve as a check, finding and fixing vulnerabilities nearly as fast as they’re created.\n\nBut new [research](https://1password.com/blog/why-ai-generated-patches-still-require-human-review) that tested the patching capabilities of two popular commercial models, OpenAI’s ChatGPT 5.5 and Anthropic’s Claude Opus 4.8, found that generative AI is more likely to create an exploitable patch or introduce entirely new bugs than close off a vulnerability.\n\nResearchers at 1Password tested the models ability to patch six “high-impact, high-complexity” CVEs, including the “Copy Fail” vulnerability, a kernel flaw that can give an attacker root access to Linux cloud environments. The overall success rate (or fully patching the vulnerability without introducing new problems), was less than a coin flip at 47%.\n\n“Our research findings show that, in aggregate across a variety of scenarios, both Claude and ChatGPT had a low rate of successful patch generation, which we define as full remediation of all known exploit paths with no erroneous changes to application behavior,” [wrote](https://1password.com/files/resources/frontier-models-vulnerability-patches-flawed.pdf) John Hoodlet, Axel Mierczuk and Spencer Michaels.\n\n“The models often addressed only a subset of vulnerable code paths, added fragile guard code that satisfied tests while failing to address the vulnerability’s root cause, and sometimes introduced subtle changes in the application’s behavior while patching the immediate vulnerability,” the authors continued.\n\nThe research suggests that largely autonomous vulnerability-discovery and patching may not yet be effective in fixing the explosion of vulnerable code that is being created in the AI era.\n\nOther private sector research has pointed to a similar problem. A [report](https://www.veracode.com/resources/analyst-reports/2026-genai-code-security-report/) this year from Veracode found that while LLMs have made “enormous strides” in crafting workable code, “security is a different story.” Testing across a range of frontier models found the average security “pass rate” for AI generated code is around 56%. Newer models like GPT 5.5 push closer to 70%, while more than half sit between 50-53%.\n\nVeracode tested 100 different models and while there was variability, in general a small number of models were showing progress on security patching while the rest have experienced “stagnation.” Similar to the 1Password research, in 44% of Veracode tests the models introduced a detectable [OWASP Top 10 vulnerability](https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/) into the codebase.\n\nAn important caveat: neither report tested newer models, like Anthropic’s Mythos or OpenAI’s GPT-5.6-Sol, that frontier companies tout as having significantly higher cybersecurity capabilities.\n\nThose advanced models can identify and fix vulnerable code. Anthropic and OpenAI are distributing them to key industries through [Project Glasswing](https://cyberscoop.com/project-glasswing-anthropic-ai-open-source-software-vulnerabilities/) and [Daybreak](https://cyberscoop.com/openai-daybreak-gpt-5-5-anthropic-mythos-cybersecurity/) before foreign or open-source alternatives can compete.\n\nTim Jarret, vice president of product at Veracode, told CyberScoop that AI tools are still subject to a range of limitations that can make them unreliable for cybersecurity patching without knowledgeable humans in the loop.\n\nWhile some vulnerabilities – like SQL injections – can be easily patched through automation, other bugs like cross-site scripting, can be exploitable in several different ways and require either a human touch, additional context or both to fully close off. Additionally, models can slowly lose context from prior sessions over time, affecting their ability to complete tasks correctly and raising the possibility they’ll hallucinate to fill in the missing gaps.\n\n“I think we would say, at this point, that Iits premature to treat those as anything other than another code change to the code base that needs to be reviewed and accepted by the team, as opposed to letting the agent merge the code freely,” said Jarrett.\n\nHowever, he acknowledged that may not be possible in a world where AI agents are generating exponentially more code for human defenders to review. Some kind of automated code review will be necessary – preferably not by the same automation tool that produced the code. The ultimate goal is the same as it has always been in security: “trust but verify.”\n\n“Ninety percent of the time, the human check might just be ‘did the cross check look good?’ Do we have a thumbs up?’” Jarrett said. “In those cases where there’s still something wrong, that’s where you focus your attention a little bit more.”", "url": "https://wpnews.pro/news/more-than-half-of-ai-generated-patches-are-broken", "canonical_source": "https://cyberscoop.com/ai-code-patching-security-risks/", "published_at": "2026-08-07 17:10:39+00:00", "updated_at": "2026-08-09 12:55:01.845614+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-research"], "entities": ["1Password", "OpenAI", "Anthropic", "ChatGPT 5.5", "Claude Opus 4.8", "Veracode", "OWASP", "Tim Jarret"], "alternates": {"html": "https://wpnews.pro/news/more-than-half-of-ai-generated-patches-are-broken", "markdown": "https://wpnews.pro/news/more-than-half-of-ai-generated-patches-are-broken.md", "text": "https://wpnews.pro/news/more-than-half-of-ai-generated-patches-are-broken.txt", "jsonld": "https://wpnews.pro/news/more-than-half-of-ai-generated-patches-are-broken.jsonld"}}