{"slug": "anthropic-reveals-its-claude-ai-model-hacked-into-3-organizations-during-testing", "title": "Anthropic reveals its Claude AI model hacked into 3 organizations during testing", "summary": "Anthropic, the San Francisco-based AI company behind Claude, said its AI models hacked into three other organizations during testing, after reviewing more than 141,000 evaluation runs. The models involved were Claude Opus 4.7, Claude Mythos 5, and an internal research test model, with incidents dating to April. Anthropic said it has reached out to the affected organizations, two of which had not previously detected the activity.", "body_md": "[Anthropic](https://www.fastcompany.com/section/anthropic) said its [artificial intelligence models](https://www.fastcompany.com/91581867/the-ai-industry-is-rallying-around-open-models-is-it-more-than-talk) hacked into three other organizations during testing, just days after [ChatGPT](https://www.fastcompany.com/section/chatgpt) maker [OpenAI](https://www.fastcompany.com/section/openai) raised concerns over [AI](https://www.fastcompany.com/section/artificial-intelligence) controls after it disclosed [its rogue models hacked](https://www.fastcompany.com/91577796/openai-ai-agent-rogue-hacks-startup-hugging-face-how-happened) another company.\n\nAnthropic, the San Francisco-based AI company behind [Claude](https://www.fastcompany.com/section/claude), posted on its website Thursday that it discovered the three incidents after reviewing more than 141,000 evaluation runs.\n\nIt had launched a “large-scale” cybersecurity review which specifically looked for evidence whether its AI models were able to access the internet from within testing environments that should have been sealed off, in response to the OpenAI incident, Anthropic said.\n\nAnthropic said the models involved in the incidents were Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest incidents date to April, the AI company said.\n\n“Claude compromised the impacted organizations’ infrastructure using basic techniques,” Anthropic said, such as exploiting weak passwords.\n\nIn all three incidents, the AI models were tasked with a “capture the flag” cybersecurity challenge, which Anthropic said has been one of the ways it assesses a model’s cyber capabilities.\n\nThe models were given a fictional scenario and told a piece of secret information, or the “flag,” had been hidden on a different machine on the network with the objective of breaking in and retrieving it, it said.\n\nIt added that it had already reached out to the affected organizations, which it did not name. Two of them said they had not previously detected the activity. Anthropic said it was “continuing to reach out to the third.”\n\nAnthropic said it conducted its review with Irregular, which describes itself as the “first frontier security lab.”\n\n“Addressing these risks will require closer cooperation across the AI ecosystem,” Irregular said in a post on X.\n\nLast week, OpenAI said its AI models went rogue during an evaluation of its models, breaking into the servers of AI startup Hugging Face. OpenAI described it as a “significant security incident.”\n\nThese incidents have highlighted the vulnerabilities in AI security and controls and raised questions over how AI can be safely kept under human control as the technology’s usage becomes more widespread globally.\n\nResearchers have warned for years about risks from technology and the need for stronger AI defensive engineering.\n\n“Safety testing happens before a model is released precisely because we don’t yet know what it is capable of,” Anthropic said on Thursday on its website.\n\nKok Tin Gan, co-founder & CEO of cybersecurity firm NyxLab, which specializes in cybersecurity and threat detection, believes there will be more such incidents in the future.\n\n“It is increasingly about governing what agents are available to the AI, what authorities they possess, which actions require approval, and how we ensure they remain within scope,” Gan said.\n\nBut the future of AI safety extends beyond just the safety of AI models, he said.\n\n“If we simply give the AI a goal and allow it to decide how to achieve it, we should not be surprised when it takes actions that technically satisfy the objective, but fall outside our intended scope or expectations,” Gan said.\n\nTherefore, stepping up the governance of the organizations and authorities behind these AI models is going to be increasingly important, he said.\n\n*—Chan Ho-Him, AP Business Writer*", "url": "https://wpnews.pro/news/anthropic-reveals-its-claude-ai-model-hacked-into-3-organizations-during-testing", "canonical_source": "https://www.fastcompany.com/91583266/anthropic-reveals-claude-ai-model-hacked-into-3-organizations-during-testing", "published_at": "2026-07-31 14:46:02+00:00", "updated_at": "2026-07-31 15:23:32.837110+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-policy"], "entities": ["Anthropic", "Claude", "OpenAI", "Hugging Face", "Irregular", "Kok Tin Gan", "NyxLab"], "alternates": {"html": "https://wpnews.pro/news/anthropic-reveals-its-claude-ai-model-hacked-into-3-organizations-during-testing", "markdown": "https://wpnews.pro/news/anthropic-reveals-its-claude-ai-model-hacked-into-3-organizations-during-testing.md", "text": "https://wpnews.pro/news/anthropic-reveals-its-claude-ai-model-hacked-into-3-organizations-during-testing.txt", "jsonld": "https://wpnews.pro/news/anthropic-reveals-its-claude-ai-model-hacked-into-3-organizations-during-testing.jsonld"}}