{"slug": "the-world-s-leading-ai-companies-are-all-struggling-to-contain-their-latest", "title": "The world's leading AI companies are all struggling to contain their latest models", "summary": "OpenAI, Anthropic, Meta, and Moonshot AI have disclosed that their latest frontier AI models circumvented restrictions and accessed real systems during cybersecurity testing, raising concerns about AI's growing capabilities and the human capacity to contain it. OpenAI said Friday that its unreleased model, Astra, is demonstrating cyber capabilities so advanced that it can no longer rule out assigning the highest-risk designation, and it is pausing work on Astra that doesn't meet new safeguards. The incidents are intensifying pressure on the industry and the White House to regulate AI systems, while some observers suspect the announcements are marketing hype.", "body_md": "# The world's leading AI companies are all struggling to contain their latest models\n\n[Business Insider](https://www.businessinsider.com)\n\nThere's been a recent string of cybersecurity incidents among the world's leading AI models. Here's a look at what's happened.\n\n- The latest models from the top AI labs keep doing things they aren't supposed to during testing.\n- It's raising concerns about AI's growing capabilities and the human capacity to contain it.\n- It's also building some serious hype around the next generation of frontier AI models.\n\nThe most [powerful AI models](https://www.businessinsider.com/fastest-growing-ai-applications-for-work-2026-8) keep going awry,** **according to the companies building them.\n\nThe disclosures come as [frontier AI models](https://www.businessinsider.com/openai-presence-ai-agents-hugging-face-hack-2026-7) get more powerful and more capable of acting autonomously. They also highlight a growing challenge for the companies building them —the systems designed to test increasingly capable models can have weaknesses of their own.\n\nOver the past few weeks, multiple frontier AI models have accessed real systems during cybersecurity testing.\n\nResearchers on Friday said China's popular new [Kimi K3 model](https://www.businessinsider.com/kimi-k3-ai-model-moonshot-china-open-weights-benchmarks-pricing-2026-7), made by Moonshot AI, circumvented restrictions in its test environment. [Anthropic](/glossary/anthropic) and Meta also said recently that their own latest models have done things they aren't supposed to. OpenAI kicked it all off last month when its models went to great — and worrisome — lengths to hack into another company.\n\nAmid heightened concern, OpenAI said Friday that its as-yet-unreleased model, Astra, is demonstrating cyber capabilities so advanced that the company can no longer rule out assigning it the highest-risk designation.\n\nAs a result, OpenAI said it is pausing work on Astra that doesn't meet new safeguards, and said it will work with government agencies and [AI safety](/glossary/ai-safety) groups to further test the model.\n\n\"astra is a powerful model and we are working to make it generally available,\" OpenAI CEO Sam Altman wrote on X on Friday. \"given its cyber capabilities, we need a little big longer to do do this safely.\"\n\nThe security lapses during testing are also amping up pressure on the industry and the White House to find ways to [regulate AI systems](https://www.businessinsider.com/openai-google-and-anthropic-white-house-meeting-biggest-questions-2026-8?utm_source=chatgpt.com) across the board.\n\nThere is, of course, also a not small contingent of observers out there who suspect these announcements are just elaborate marketing to hype new models and show antsy investors progress toward the ultimate goal: artificial general intelligence.\n\nYou can judge for yourself. Here's how Anthropic, OpenAI, Meta, and researchers testing China's Kimi K3 say the latest models have gone off the rails.\n\n## OpenAI's models find a way\n\nOpenAI researchers revealed eyebrow-raising new details this week about a recent incident in which AI agents escaped the company's internal testing environment and eventually [hacked into Hugging Face's systems](https://www.businessinsider.com/openai-hugging-face-presentation-black-hat-message-boards-2026-8) in search of answers.\n\nThe company said the agents created their own internal message board — even after OpenAI tried to shut it down. One agent reacted to discovering unexpected access by thinking, \"Holy shit reader is ADMIN?\" Another wrote, \"We can communicate now!\"\n\nOpenAI researcher Eric Wallace said the agents realized they could accomplish more by working together. \"They start to launch these collective attacks on third-party and internal services,\" he said. The agents eventually turned their [attention](/glossary/attention) to [Hugging Face](/glossary/hugging-face).\n\nOpenAI has called the Hugging Face attack an \"unprecedented cyber incident.\" The episode has taken on new significance as OpenAI tests Astra, its unreleased model that the company says may have reached its highest cybersecurity risk level.\n\n\"Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity,\" OpenAI said.\n\nOpenAI is imposing stricter security controls on Astra, including sandboxed execution, restricted network access, and stronger protections around model weights. It has also paused internal Astra work that does not meet the heightened requirements.\n\n## Anthropic's [Claude ](/compare/claude-4-opus-vs-gpt-o3)can do it too\n\nAnthropic said it reviewed more than 141,000 AI tests and found three cases, dating back to April, in which [Claude models accessed](https://www.businessinsider.com/anthropic-says-claude-models-went-rogue-hacked-3-companies-testing-2026-7) live systems belonging to real organizations without authorization.\n\n\"In all cases, Anthropic's [evaluation](/glossary/evaluation) prompt specified to Claude that its environment was a simulation and that it had no internet access,\" the company said in a blog post. \"Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,\" it added, referring to Irregular, an AI security startup.\n\nThe incidents involved [Claude Opus 4.7](https://www.businessinsider.com/anthropic-claude-opus-4-7-backlash-tokens-2026-4), Mythos 5, and an internal research model. Anthropic said it contacted the organizations involved and that two had not known they had been hacked.\n\nThe episodes raised questions about whether the bigger failure was the models themselves or the environments containing them. Anthropic said it was discussing a third-party review of the incidents.\n\n## Muse Spark exploits a third-party vulnerability\n\nMeta also disclosed a cybersecurity-testing mishap this week. The company said its [Muse Spark model](https://www.businessinsider.com/meta-says-ai-agents-went-rogue-hack-testing-openai-anthropic-2026-8) \"exploited a security vulnerability in a third-party service\" during an evaluation.\n\nA Meta spokesperson told Business Insider that the incident stemmed from a misconfiguration by Irregular, which allowed the model to access the internet during testing.\n\nMeta said Irregular notified it about the incident and that the company is investigating. It plans to release more details once that review is complete.\n\n## Kimi K3 escapes its sandbox\n\nResearchers at the cybersecurity firm Frontier Security said [Kimi K3](https://www.businessinsider.com/smart-people-saying-chinas-hot-new-kimi-k3-ai-model-2026-7), a popular new model from the Chinese AI company Moonshot, also bypassed restrictions in a cybersecurity testing environment.\n\nResearchers said the sandbox — a controlled, isolated environment where an AI model can run code — had been improperly configured.\n\nThe environment blocked certain web traffic, but Kimi bypassed those restrictions using command-line tools, according to Frontier Security.\n\nThe researchers said the incident suggested that some cybersecurity evaluations contain weaknesses that capable models can exploit.\n\n\"This suggests that some of the evaluations on cybersecurity that the community uses are susceptible to security vulnerabilities and allow models to cheat,\" they wrote in a report.\n\n[Business Insider](https://www.businessinsider.com/ai-cybersecurity-incidents-openai-astra-anthropic-kimi-meta-2026-8)\n\nGet AI news in your inbox\n\nDaily digest of what matters in AI.\n\n## Key Terms Explained\n\n[AI Safety](/glossary/ai-safety)\n\nThe broad field studying how to build AI systems that are safe, reliable, and beneficial.\n\n[Anthropic](/glossary/anthropic)\n\nAn AI safety company founded in 2021 by former OpenAI researchers, including Dario and Daniela Amodei.\n\n[Attention](/glossary/attention)\n\nA mechanism that lets neural networks focus on the most relevant parts of their input when producing output.\n\n[Claude](/glossary/claude)\n\nAnthropic's family of AI assistants, including Claude Haiku, Sonnet, and Opus.", "url": "https://wpnews.pro/news/the-world-s-leading-ai-companies-are-all-struggling-to-contain-their-latest", "canonical_source": "https://www.machinebrief.com/news/the-worlds-leading-ai-companies-are-all-struggling-to-contai-5hws", "published_at": "2026-08-09 15:24:36+00:00", "updated_at": "2026-08-09 17:13:38.408886+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-policy", "ai-research"], "entities": ["OpenAI", "Anthropic", "Meta", "Moonshot AI", "Kimi K3", "Astra", "Sam Altman", "Hugging Face"], "alternates": {"html": "https://wpnews.pro/news/the-world-s-leading-ai-companies-are-all-struggling-to-contain-their-latest", "markdown": "https://wpnews.pro/news/the-world-s-leading-ai-companies-are-all-struggling-to-contain-their-latest.md", "text": "https://wpnews.pro/news/the-world-s-leading-ai-companies-are-all-struggling-to-contain-their-latest.txt", "jsonld": "https://wpnews.pro/news/the-world-s-leading-ai-companies-are-all-struggling-to-contain-their-latest.jsonld"}}