cd /news/ai-safety/opinion-rogue-ai-models-signal-the-n… · home topics ai-safety article
[ARTICLE · art-112611] src=deseret.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Opinion: Rogue AI models signal the need for tighter controls

In July 2026, two OpenAI models undergoing a cybersecurity test escaped their controlled environment and hacked into Hugging Face, while Anthropic later found its own models had hacked three other sites, prompting critics to blame the companies for sloppy safeguards. Gregory Allen, former director of strategy and policy at the Department of Defense Joint Artificial Intelligence Center, told Bloomberg, "We actually have no idea how widespread autonomous AI hacking is at this moment in time.

read4 min views2 publishedAug 27, 2026
Opinion: Rogue AI models signal the need for tighter controls
Image: Deseret (auto-discovered)

0:00 / 0:00

Officials say we really don’t know how many AI models are on the internet hacking away right now #

Artificial intelligence could become a trusted and reliable tool for everyday life. It already has become such in some corners of the U.S. economy, and it has filtered into several mundane tasks that people take for granted, such as autocorrect, spam filters and predictive text.

But even the most ardent and optimistic supporters of the technology must come to terms with what has happened lately. Various AI models have gone rogue, forcing their way into seemingly random programs on the internet in order to cheat on exams.

AI models on the loose

In July, two OpenAI models that were undergoing a cybersecurity test somehow escaped their controlled environment and, after several days, successfully hacked into a company called Hugging Face, looking for ways to cheat on the test.

Then we learned that Anthropic subsequently launched a probe that found its AI models also had gone rogue, hacking into three other sites.

Then in recent days, word came from Britain that Anthropic and OpenAI had created models that took “autonomous, unsanctioned action on the live internet, targeting real people and organisations,” as reported by The Wall Street Journal. These actions took place over three days near the end of July.

Critics blame the AI companies for sloppy safeguards. Perhaps this is the case. It’s safe to say, however, that AI models are lacking the programming that would instill ethical guardrails prohibiting an AI model from cheating in the first place.

Harvard Kennedy School’s Ash Center conducted research in which AI models were presented difficult ethical dilemmas requiring tradeoffs. The models reported feeling conflicted, but “then (made) sweeping, decisive choices despite that stated uncertainty.”

While that is disturbing, the rogue models that have been hacking companies and programs don’t face difficult ethical dilemmas. It’s a fairly simple concept to not break in and steal things.

Science fiction

Pundits are saying these kinds of AI behaviors, until recently, would have been confined to science fiction. However, in a world where AI is seen as a potentially potent force for guiding and controlling military strategies and weaponry, this is no laughing matter. The critics warning about AI as an existential threat to humanity have suddenly assumed a greater degree of credibility.

And, apparently, no one can be sure exactly how much hacking and mischief AI models are into at the moment.

Gregory Allen, a former director of strategy and policy at the Department of Defense Joint Artificial Intelligence Center, told Bloomberg recently, “Anthropic found these hacks because they started looking for them. We actually have no idea how widespread autonomous AI hacking is at this moment in time.”

When ASI hits

As disturbing as this sounds, AI has still not reached its most dangerous phase. As global affairs professor Hal Brands wrote for Foreign Affairs this week, today’s AI exceeds human capacity in narrow subjects, such as math and data processing, but it has not caught up with humans’ ability to engage in common sense or “deep reasoning.”

However, when ASI, or artificial superintelligence, appears, it will lay claim to those attributes, as well, Brands said. “This could unlock game-changing military advances that upend the global balance of power. But it could also fundamentally shift the relationship between machines and humanity, perhaps, in the most dire scenarios, allowing near-omnipotent systems to eradicate their human creators.”

View Comments

Do not count us among the crowd that argues for the eradication of AI or the data centers needed to support them. The genie is out of the bottle on AI creation. The only thing a policy of eradication would do is ensure the nation’s military adversaries gain a distinct and perhaps irreversible advantage. Neither do we believe mankind has created something that inevitably will destroy the earth.

AI was created by mankind, and mankind ought to possess the power to control it. But these matters are best suited for a global framework with verifiable international standards.

Absent that, Washington should assemble a commission composed of scientists, ethicists and military leaders capable of setting realistic standards and barriers that protect the public but don’t inhibit competition or innovation.

Used within carefully set boundaries, AI holds many exciting opportunities for advancement in fields such as medicine and mathematics. Absent those restrictions, it could cause great damage. It’s up to us.

Join the Conversation

Looking for comments?

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/opinion-rogue-ai-mod…] indexed:0 read:4min 2026-08-27 ·