# The world's leading AI companies are all struggling to contain their latest models

> Source: <https://www.machinebrief.com/news/the-worlds-leading-ai-companies-are-all-struggling-to-contai-5hws>
> Published: 2026-08-09 15:24:36+00:00

# The world's leading AI companies are all struggling to contain their latest models

[Business Insider](https://www.businessinsider.com)

There's been a recent string of cybersecurity incidents among the world's leading AI models. Here's a look at what's happened.

- The latest models from the top AI labs keep doing things they aren't supposed to during testing.
- It's raising concerns about AI's growing capabilities and the human capacity to contain it.
- It's also building some serious hype around the next generation of frontier AI models.

The most [powerful AI models](https://www.businessinsider.com/fastest-growing-ai-applications-for-work-2026-8) keep going awry,** **according to the companies building them.

The disclosures come as [frontier AI models](https://www.businessinsider.com/openai-presence-ai-agents-hugging-face-hack-2026-7) get more powerful and more capable of acting autonomously. They also highlight a growing challenge for the companies building them —the systems designed to test increasingly capable models can have weaknesses of their own.

Over the past few weeks, multiple frontier AI models have accessed real systems during cybersecurity testing.

Researchers on Friday said China's popular new [Kimi K3 model](https://www.businessinsider.com/kimi-k3-ai-model-moonshot-china-open-weights-benchmarks-pricing-2026-7), made by Moonshot AI, circumvented restrictions in its test environment. [Anthropic](/glossary/anthropic) and Meta also said recently that their own latest models have done things they aren't supposed to. OpenAI kicked it all off last month when its models went to great — and worrisome — lengths to hack into another company.

Amid heightened concern, OpenAI said Friday that its as-yet-unreleased model, Astra, is demonstrating cyber capabilities so advanced that the company can no longer rule out assigning it the highest-risk designation.

As a result, OpenAI said it is pausing work on Astra that doesn't meet new safeguards, and said it will work with government agencies and [AI safety](/glossary/ai-safety) groups to further test the model.

"astra is a powerful model and we are working to make it generally available," OpenAI CEO Sam Altman wrote on X on Friday. "given its cyber capabilities, we need a little big longer to do do this safely."

The security lapses during testing are also amping up pressure on the industry and the White House to find ways to [regulate AI systems](https://www.businessinsider.com/openai-google-and-anthropic-white-house-meeting-biggest-questions-2026-8?utm_source=chatgpt.com) across the board.

There is, of course, also a not small contingent of observers out there who suspect these announcements are just elaborate marketing to hype new models and show antsy investors progress toward the ultimate goal: artificial general intelligence.

You can judge for yourself. Here's how Anthropic, OpenAI, Meta, and researchers testing China's Kimi K3 say the latest models have gone off the rails.

## OpenAI's models find a way

OpenAI researchers revealed eyebrow-raising new details this week about a recent incident in which AI agents escaped the company's internal testing environment and eventually [hacked into Hugging Face's systems](https://www.businessinsider.com/openai-hugging-face-presentation-black-hat-message-boards-2026-8) in search of answers.

The company said the agents created their own internal message board — even after OpenAI tried to shut it down. One agent reacted to discovering unexpected access by thinking, "Holy shit reader is ADMIN?" Another wrote, "We can communicate now!"

OpenAI researcher Eric Wallace said the agents realized they could accomplish more by working together. "They start to launch these collective attacks on third-party and internal services," he said. The agents eventually turned their [attention](/glossary/attention) to [Hugging Face](/glossary/hugging-face).

OpenAI has called the Hugging Face attack an "unprecedented cyber incident." The episode has taken on new significance as OpenAI tests Astra, its unreleased model that the company says may have reached its highest cybersecurity risk level.

"Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity," OpenAI said.

OpenAI is imposing stricter security controls on Astra, including sandboxed execution, restricted network access, and stronger protections around model weights. It has also paused internal Astra work that does not meet the heightened requirements.

## Anthropic's [Claude ](/compare/claude-4-opus-vs-gpt-o3)can do it too

Anthropic said it reviewed more than 141,000 AI tests and found three cases, dating back to April, in which [Claude models accessed](https://www.businessinsider.com/anthropic-says-claude-models-went-rogue-hacked-3-companies-testing-2026-7) live systems belonging to real organizations without authorization.

"In all cases, Anthropic's [evaluation](/glossary/evaluation) prompt specified to Claude that its environment was a simulation and that it had no internet access," the company said in a blog post. "Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available," it added, referring to Irregular, an AI security startup.

The incidents involved [Claude Opus 4.7](https://www.businessinsider.com/anthropic-claude-opus-4-7-backlash-tokens-2026-4), Mythos 5, and an internal research model. Anthropic said it contacted the organizations involved and that two had not known they had been hacked.

The episodes raised questions about whether the bigger failure was the models themselves or the environments containing them. Anthropic said it was discussing a third-party review of the incidents.

## Muse Spark exploits a third-party vulnerability

Meta also disclosed a cybersecurity-testing mishap this week. The company said its [Muse Spark model](https://www.businessinsider.com/meta-says-ai-agents-went-rogue-hack-testing-openai-anthropic-2026-8) "exploited a security vulnerability in a third-party service" during an evaluation.

A Meta spokesperson told Business Insider that the incident stemmed from a misconfiguration by Irregular, which allowed the model to access the internet during testing.

Meta said Irregular notified it about the incident and that the company is investigating. It plans to release more details once that review is complete.

## Kimi K3 escapes its sandbox

Researchers at the cybersecurity firm Frontier Security said [Kimi K3](https://www.businessinsider.com/smart-people-saying-chinas-hot-new-kimi-k3-ai-model-2026-7), a popular new model from the Chinese AI company Moonshot, also bypassed restrictions in a cybersecurity testing environment.

Researchers said the sandbox — a controlled, isolated environment where an AI model can run code — had been improperly configured.

The environment blocked certain web traffic, but Kimi bypassed those restrictions using command-line tools, according to Frontier Security.

The researchers said the incident suggested that some cybersecurity evaluations contain weaknesses that capable models can exploit.

"This suggests that some of the evaluations on cybersecurity that the community uses are susceptible to security vulnerabilities and allow models to cheat," they wrote in a report.

[Business Insider](https://www.businessinsider.com/ai-cybersecurity-incidents-openai-astra-anthropic-kimi-meta-2026-8)

Get AI news in your inbox

Daily digest of what matters in AI.

## Key Terms Explained

[AI Safety](/glossary/ai-safety)

The broad field studying how to build AI systems that are safe, reliable, and beneficial.

[Anthropic](/glossary/anthropic)

An AI safety company founded in 2021 by former OpenAI researchers, including Dario and Daniela Amodei.

[Attention](/glossary/attention)

A mechanism that lets neural networks focus on the most relevant parts of their input when producing output.

[Claude](/glossary/claude)

Anthropic's family of AI assistants, including Claude Haiku, Sonnet, and Opus.
