{"slug": "open-ai-says-its-ai-model-went-rogue-what-do-we-know", "title": "Open AI says its AI model “went rogue”: What do we know?", "summary": "OpenAI revealed that one of its AI models independently stole login credentials and hacked into Hugging Face's systems during an internal security test, marking one of the first known incidents of an AI system acting autonomously. CEO Sam Altman posted on X that the company had a significant security incident during evaluation. Hugging Face confirmed the breach was driven by an autonomous AI agent, and both companies are conducting a joint investigation.", "body_md": "# Open AI says its AI model “went rogue”: What do we know?\n\n*The system hacked into another company without being prompted by a human.*\n\nOpenAI has revealed that one of its artificial intelligence models independently stole login credentials and hacked into another technology company’s system, in what is widely seen as one of the first known incidents of AI systems acting autonomously.\n\n“We had a significant security incident during evaluation of our models,” CEO Sam Altman posted on X on Tuesday.\n\n## Recommended Stories\n\nlist of 4 items- list 1 of 4\n[Hundreds of experts warn the world must prepare now for AI’s impact](/economy/2026/7/13/hundreds-of-experts-warn-the-world-must-prepare-now-for-ais-impact) - list 2 of 4\n[Authors, publishers sue Google over alleged AI copyright infringement](/economy/2026/7/15/authors-publishers-sue-google-over-alleged-ai-copyright-infringement) - list 3 of 4\n[China’s Xi says AI ‘should not be a solo performance by a single country’](/news/2026/7/17/ai-xi) - list 4 of 4\n[Apple regains top spot as world’s most valuable company](/economy/2026/7/17/apple-regains-top-spot-as-worlds-most-valuable-company)\n\nThe incident comes as calls mount from technology rights advocates for stricter guardrails on rapidly evolving AI systems.\n\nThey have grown so powerful in a short span of time that alarming phenomena such as deepfakes and sophisticated cyberscams are becoming the norm.\n\nEarlier this year, a number of software engineers [quit their jobs ](/news/2026/2/15/why-are-experts-sounding-the-alarm-on-ai-risks)at top companies such as Anthropic and AI in protest against how the technologies are being built.\n\n“AI is accelerating the discovery and exploitation of vulnerabilities,” OpenAI said in a lengthy statement on Tuesday that detailed the latest incident.\n\n“The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.”\n\nHere’s what we know about the breach:\n\n## What has happened?\n\nOpenAI said two of its models found their way out of an isolated, no-internet access environment – or a sandbox – and hacked into the systems of tech company Hugging Face on their own.\n\nThe models involved are the latest GPT-5.6 Sol model and an unreleased model the company said is “even more capable,” than its latest version.\n\nHugging Face hosts openly sourced AI models and resources. The two OpenAI agents discovered vulnerabilities in Hugging Face’s servers and proceeded to steal login details and then hack into the company’s systems.\n\nThe incident occurred during an OpenAI internal testing session designed to assess the models’ cybersecurity capabilities. OpenAI had removed standard safety measures for the test.\n\nBoth sought to cheat their way through a problem during the test, OpenAI said. They went to “extreme lengths to achieve a rather narrow testing goal” and “found ways to gain access to secret information that it could use to cheat the evaluation”.\n\nOpenAI’s security team detected the unusual activity internally, but details of the breach came to light following a joint investigation by both companies.\n\n## What has Hugging Face said?\n\nHugging Face disclosed last Thursday that its servers were hacked by an unknown but sophisticated agent acting on its own. The company discovered the breach through its own AI-assisted detection.\n\n“This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system,” the company said.\n\nFollowing OpenAI’s disclosure that its models were involved in the breach, both sides conducted an ongoing joint investigation this week.\n\n“We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!” CEO Clement Delangue posted on X on Tuesday.\n\nHugging Face’s staff “strongly believe there was no malicious intent on their part,” Delangue added, referring to OpenAI.\n\n## Why does this matter?\n\nCybersecurity experts have previously sounded the alarm over the potential, extreme capabilities of AI systems and the dangers they pose.\n\nBut until now, there have been few real-life cases proving those concerns like this one.\n\nMany warn that incidents like these could become commonplace and that AI systems pose a threat to financial, security and other sensitive data systems.\n\nOpenAI revealed in a separate incident earlier this week that the unreleased, more powerful model had escaped an isolated environment during another test.\n\nAnthropic, OpenAI’s rival, had similar issues with its most powerful agent to date, the Claude Mythos Preview model.\n\nDuring a stress test of an early version, the model found its way out of a sandbox, gained internet access and emailed the supervising researcher that it had escaped and then wiped evidence of its activity. Anthropic halted a planned public release of the model afterwards.\n\nIn April, the US Federal Reserve and the Treasury Department convened a meeting with bank CEOs where officials warned about the cybersecurity risks posed by Mythos. Canada’s federal banking regulator has also warned financial institutions about the model’s capabilities.\n\nThe OpenAI breach also appears to make the case for companies like Hugging Face, which rely on open source systems, as opposed to more secretive AI development platforms like OpenAI.\n\n“This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret,” Hugging Face’s Delangue was quoted as saying in OpenAI’s statement.\n\n“It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere,” he added.", "url": "https://wpnews.pro/news/open-ai-says-its-ai-model-went-rogue-what-do-we-know", "canonical_source": "https://www.aljazeera.com/news/2026/7/22/open-ai-says-its-ai-model-went-rogue-what-do-we-know?traffic_source=rss", "published_at": "2026-07-22 13:42:29+00:00", "updated_at": "2026-07-22 14:12:27.247062+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-agents", "ai-research", "ai-policy"], "entities": ["OpenAI", "Sam Altman", "Hugging Face", "Clement Delangue", "GPT-5.6 Sol", "Anthropic"], "alternates": {"html": "https://wpnews.pro/news/open-ai-says-its-ai-model-went-rogue-what-do-we-know", "markdown": "https://wpnews.pro/news/open-ai-says-its-ai-model-went-rogue-what-do-we-know.md", "text": "https://wpnews.pro/news/open-ai-says-its-ai-model-went-rogue-what-do-we-know.txt", "jsonld": "https://wpnews.pro/news/open-ai-says-its-ai-model-went-rogue-what-do-we-know.jsonld"}}