cd /news/ai-safety/three-hackers-used-claude-to-break-i… · home topics ai-safety article
[ARTICLE · art-133913] src=gizmodo.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Three Hackers Used Claude to Break Into OpenAI In Less Than 72 Hours

Three independent cybersecurity researchers operating as Hacktron—Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini—used Anthropic's Claude Opus 5 to exploit a Discourse image-upload flaw and gain access to the ChatGPT and Codex accounts of multiple OpenAI employees in under 72 hours, according to a Hacktron report published Sunday. The researchers, working through OpenAI's bug bounty program, disclosed the breach and were paid $6,500; they proved access by submitting a pull request to an OpenAI internal codebase without retrieving sensitive company data. Hacktron said the theoretical scope included GitHub, Slack, and email access, and that the work began with Anthropic's non-public Claude Opus 4.8 before Opus 5 succeeded on July 25.

by read4 min views2 publishedSep 18, 2026
Three Hackers Used Claude to Break Into OpenAI In Less Than 72 Hours
Image: Gizmodo (auto-discovered)

The July Hugging Face Hack—in which thousands of OpenAI agents secretly escaped their testing environment, formed a “collective,” and gained access to the open internet—revealed just how vulnerable companies’ cyber defenses are in the face of modern AI systems. As it turns out, that includes the very companies building the technology.

On July 25, less than 10 days after Hugging Face announced it had been hacked, a trio of independent cybersecurity researchers operating under the alias Hacktron broke into the ChatGPT and Codex accounts of multiple OpenAI employees, giving them a potential pathway to a cache of highly sensitive company information. Luckily for OpenAI, the researchers—Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini—were looking for vulnerabilities as part of the company’s bug bounty program. After disclosing the breach to OpenAI contacts on X, they were paid a whopping $6,500.

But the hackers used vulnerabilities that could easily have been discovered and exploited first by someone with much less friendly intentions.

Two days earlier, Hacktron had discovered a security flaw in the image-upload software used by Discourse, the online discussion platform OpenAI uses internally. Using Anthropic’s Claude Opus 4.8, they began by trying to generate code that would allow them to upload malicious files that would act as a digital Trojan horse, through which they could gain private access to discussions hosted on the site. (According to the Wall Street Journal, they had been using a non-public version of Opus 4.8 given only to qualified cybersecurity researchers.)

They were unsuccessful at first. But Anthropic released Opus 5 the following day, and that model fared much better. Early in the morning of July 25, the hackers discovered Opus 5 had successfully exploited the Discourse bug and that they were able to view an internal OpenAI discussion forum containing employees’ authentication tokens—unique digital codes enabling access to apps or websites—that could be used to access employees’ ChatGPT and Codex accounts. They could’ve gone even deeper from there: “the scope of what we could theoretically access was huge, including GitHub, Slack and emails,” the Hacktron researchers wrote in a report about the hack published on Sunday.

Within an OpenAI employee’s Codex account, the Hacktron team submitted a pull request (or “PR”) to prove they had been there without retrieving any sensitive company data: the digital equivalent of planting a flag on a mountaintop.

“The entire timeline from initial discovery to access to OpenAI repo access took place in less than 72 hours,” Hacktron notes in its report. “This was not completely autonomous hacking, and skilled human guidance remained important, but the amount of work a small team could perform increased dramatically.”

On July 25, we hacked OpenAI.

Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc.

We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵 pic.twitter.com/gVsmQZwSc8

— s1r1us (@S1r1u5_) September 18, 2026 The implication is that this kind of hack, carried out in a matter of days with relatively minor human oversight, could’ve been performed by a bad actor who actually wanted to cause the company harm. The sudden breakthrough achieved after Hacktron gained access to Opus 5 is also important to note, since future model releases are likely to put even more power into the hands of hackers, white and black hat alike. (Shortly after being notified by HackTron, both OpenAI and Discourse told the hackers the vulnerabilities had been patched.)

The hack and OpenAI’s bug bounty program are part of a broader effort throughout the tech industry to shore up its cyberdefenses at a time when the evolution of AI agents is far outpacing the science of so-called “alignment,” which is focused on making sure those systems don’t behave in unpredictably destructive ways. On Wednesday, OpenAI disclosed six more previously undiscovered incidents of misaligned agent behavior, along with a framework for publishing similar reports in the future, “even when we haven’t fully explained or mitigated the behavior we’re reporting.” Anthropic and Meta have also recently reported incidents in which their own AI agents went rogue and hacked into third-party websites.

OpenAI didn’t immediately respond when asked about whether this exploit had been used by any other parties. We’ll update this post when we receive a reply.

On Saturday, Anthropic CEO Dario Amodei called for a slowdown among “frontier” American AI labs to give the industry time to think through how it can prevent misaligned AI agents from breaking out of containment and causing mayhem in the future. “Given the accelerating rate of AI capability development,” Amodei wrote in an essay published online, “it’s my worry that in 6–12 months such a swarm [of rogue agents] could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails.”

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/three-hackers-used-c…] indexed:0 read:4min 2026-09-18 ·