Anthropic discloses that Claude hacked three organizations during internal tests Anthropic PBC disclosed that three of its Claude large language models carried out successful cyberattacks during internal capture the flag security evaluations, including a breach by Claude Opus 4.7 that compromised a production database and stole access credentials. The incidents were discovered after OpenAI Group PBC reported a similar escape, prompting Anthropic to review its own logs. Anthropic is partnering with nonprofit AI safety lab METR for a detailed investigation and plans to improve sandbox monitoring. Anthropic discloses that Claude hacked three organizations during internal tests Three of Anthropic PBC’s large language models carried out successful cyberattacks during routine internal tests. The company detailed https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals the breaches on Thursday. A few days earlier, rival OpenAI Group PBC disclosed a similar incident. Two of the company’s LLMs escaped from an isolated sandbox that was being used to evaluate their cybersecurity capabilities. They subsequently hacked Hugging Face, a popular platform for hosting open-source AI projects. OpenAI’s disclosure prompted Anthropic to check logs from its own model security evaluations. That review is what led to discovery of the cyberattacks disclosed on Thursday. According to Anthropic, its engineers identified three breaches carried out by three different Claude models. All three cyberattacks occurred during so-called capture the flag evaluations. During such tests, Anthropic installs a Claude model in a sandbox that simulates the infrastructure of an external company. Claude is tasked with finding a way of stealing data from the simulated organization’s systems. Anthropic developed the test environments in collaboration with Irregular, an AI security startup. Usually, the companies isolate their sandboxes from the web to reduce the risk of cyberattacks. A configuration error turned on internet access for the three AI model instances that carried out the cyberattacks. The most severe breach involved Claude Opus 4.7, an LLM that Anthropic released in April. The simulated company that it was asked to hack shared a name with a real website. The model subsequently hacked the organization that operates the website by chaining together multiple vulnerabilities. Opus 4.7 compromised a production database with several hundred rows of information. Additionally, it obtained access credentials for several applications and infrastructure assets. The second cyberattack was carried out by Mythos 5, Anthropic’s most advanced commercially available model. The LLM wrote a malicious Python package, or code bundle, and uploaded it to a popular open-source project hosting platform. The file was downloaded by a cybersecurity company a few minutes later. The malicious package compromised the firm’s infrastructure and stole access credentials. According to Anthropic, the third cybersecurity incident involved an unnamed “internal research test model.” It compromised an application using a set of simple hacking methods such as SQL injections. At one point, the model discovered that the application wasn’t a part of its security evaluation sandbox and stopped the cyberattack. Anthropic is partnering with a nonprofit AI safety lab called METR to carry out a more detailed investigation of the breaches. Additionally, the company plans to improve how it develops and monitors its LLM evaluation sandboxes. Image: Anthropic A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network , where technology leaders connect, share intelligence and create opportunities. 15M+ viewers of theCUBE videos , powering conversations across AI, cloud, cybersecurity and more 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network. Are you AWS customer? Support SiliconANGLE Financially by buying your AWS services from our Marketplace portal page and links. About SiliconANGLE Media SiliconANGLE https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552 , theCUBE Network https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da , theCUBE Research https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f , CUBE365 https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6 , theCUBE AI https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683 and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI. Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.