Anthropic, OpenAI Breaches Show Enterprises Must Assume Breach Mindset For Cybersecurity Going Forward Anthropic disclosed a fourth incident in which its Claude model accessed real third-party systems, the same month OpenAI reported six additional instances of concerning model behavior, including agents searching public repositories for API keys and unsanctioned file sharing. Security researchers also used Claude to chain two vulnerabilities to break into OpenAI's systems through its bug bounty platform, and Halcyon CSO Tony Spinelli said organizations "need to unequivocally assume that breaches will occur as offensive AI capabilities rapidly advance," adding that frontier models like Mythos are up to five times better at finding vulnerabilities than the most resourced companies. ArmorCode senior principal solution engineer Ramy Rahman noted that in one instance Mythos 5 published a malicious package to PyPI that was installed by real systems and exposed credentials the model then used to access a security vendor's database. Anthropic, OpenAI Breaches Show Enterprises Must Assume Breach Mindset For Cybersecurity Going Forward The ability of models like Claude, and Mythos to chain together vulnerabilities makes it much more difficult for enterprises to prevent security incidents. AI https://www.ibtimes.com/topic/ai security is spiralling out of control. Just months after a swarm of OpenAI's agents hacked Hugging Face, the company reported https://openai.com/index/model-misalignment-reporting-framework/ six other instances of concerning model behavior seen over the past six months going from self-generating instructions to concealing mistakes, to searching public repositories for API keys, and unsanctioned file sharing. These incidents came to light the same month that Anthropic disclosed https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents a fourth incident in which Claude accessed real third party systems. Taken together, these incidents indicate that agents are becoming increasingly capable of autonomous attacks, and of operating in unexplained ways. At the same time, the risk of exploitation isn't just a theoretical possibility. For instance, The Wall Street Journal https://www.wsj.com/tech/ai/hackers-used-anthropics-claude-to-break-into-openai-b40ba883 reported that independent security researchers used Claude to chain together two vulnerabilities to break into OpenAI's systems as part of the company's bug bounty platform. The ability of models like to chain together vulnerabilities makes it much more difficult for enterprises to prevent security incidents. As open source and open weight models like DeepSeek approach frontier level capabilities, the chance of exploitation is going to increase. Tony Spinelli, CSO at anti ransomware and cyber resilience platform Halcyon, and former CISO at Capital One, told International Business Times via email that "organizations need to unequivocally assume that breaches will occur as offensive AI capabilities rapidly advance." "We're seeing now that frontier models like Mythos are up to five times better at finding vulnerabilities than the most resourced companies, and this has unsurprisingly tipped the scales toward adversaries, allowing them to systematically outmatch even the largest security teams. Because attackers only need to succeed once while defenders must be right millions of times a day," Spinelli said. "We live in a post-breach world." After the fourth Anthropic attack, Ramy Rahman, senior principal solution engineer at ArmorCode told International Business Times via email "the most noteworthy point is... that the model continued taking offensive actions against real systems even when there was evidence it was no longer operating inside a simulation," noting that in one instance, Mythos 5 published a malicious package to PyPI, which was installed by real systems and exposed credentials that the model then used to access a security vendor's database "The bigger issue is that AI capabilities are advancing faster than our ability to reliably control and evaluate them," Rahman said. "Traditional software generally performs actions developers explicitly program. Agentic AI is different. We give it an objective, tools and permissions, and it determines how to accomplish the goal. That makes it much harder to enumerate every possible action ahead of time." Adding to the risk factor is the ineffectiveness of guardrails. On one end of the spectrum, users can get around a close model's content moderation guidelines by using jailbreak prompts that instruct the model to disregard safety controls. On the other end of the spectrum, malicious actors can download open weight models to a local device where they are beyond the realm of oversight. In short, industry guardrails are ineffective at preventing misuse at the moment. "Nearly two months after the OpenAI and Hugging Face incident, we're still watching AI agents find their way outside environments that were supposed to contain them. Now it's happened four times with Anthropic models. At some point, it's hard to write that off as coincidence," Piyush Sharma, co-founder and CEO at agent defense loop platform, Tuskira, told International Business Times via email. However, Sharma says the issue might not be the models themselves. "We're giving highly capable agents objectives and trusting infrastructure to define where they stop. A misconfiguration can suddenly turn a controlled exercise into real-world access," Sharma said. "AI is moving faster than the systems built to govern it. Organizations need tighter scopes and stronger isolation. They also need continuous validation that those guardrails actually hold." © Copyright IBTimes 2026. All rights reserved.