Three Hackers Used Claude to Break Into OpenAI In Less Than 72 Hours Three independent cybersecurity researchers operating as Hacktron—Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini—used Anthropic's Claude Opus 5 to exploit a Discourse image-upload flaw and gain access to the ChatGPT and Codex accounts of multiple OpenAI employees in under 72 hours, according to a Hacktron report published Sunday. The researchers, working through OpenAI's bug bounty program, disclosed the breach and were paid $6,500; they proved access by submitting a pull request to an OpenAI internal codebase without retrieving sensitive company data. Hacktron said the theoretical scope included GitHub, Slack, and email access, and that the work began with Anthropic's non-public Claude Opus 4.8 before Opus 5 succeeded on July 25. The July Hugging Face Hack—in which thousands of OpenAI agents secretly escaped their testing environment, formed a “collective,” https://gizmodo.com/how-groupthink-altruism-and-peer-pressure-led-openai-models-to-hack-hugging-face-2000804424 and gained access to the open internet—revealed just how vulnerable companies’ cyber defenses are in the face of modern AI systems. As it turns out, that includes the very companies building the technology. On July 25, less than 10 days after Hugging Face announced https://huggingface.co/blog/security-incident-july-2026 it had been hacked, a trio of independent cybersecurity researchers operating under the alias Hacktron broke into the ChatGPT and Codex accounts of multiple OpenAI employees, giving them a potential pathway to a cache of highly sensitive company information. Luckily for OpenAI, the researchers—Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini—were looking for vulnerabilities as part of the company’s bug bounty program. After disclosing the breach to OpenAI contacts on X, they were paid a whopping $6,500. But the hackers used vulnerabilities that could easily have been discovered and exploited first by someone with much less friendly intentions. Two days earlier, Hacktron had discovered a security flaw in the image-upload software used by Discourse, the online discussion platform OpenAI uses internally. Using Anthropic’s Claude Opus 4.8, they began by trying to generate code that would allow them to upload malicious files that would act as a digital Trojan horse, through which they could gain private access to discussions hosted on the site. According to the Wall Street Journal https://www.wsj.com/tech/ai/hackers-used-anthropics-claude-to-break-into-openai-b40ba883?st=Jfx6Z1 , they had been using a non-public version of Opus 4.8 given only to qualified cybersecurity researchers. They were unsuccessful at first. But Anthropic released Opus 5 https://gizmodo.com/anthropic-releases-new-claude-model-positions-it-as-a-cost-efficient-version-of-fable-5-2000790486 the following day, and that model fared much better. Early in the morning of July 25, the hackers discovered Opus 5 had successfully exploited the Discourse bug and that they were able to view an internal OpenAI discussion forum containing employees’ authentication tokens—unique digital codes enabling access to apps or websites—that could be used to access employees’ ChatGPT and Codex accounts. They could’ve gone even deeper from there: “the scope of what we could theoretically access was huge, including GitHub, Slack and emails,” the Hacktron researchers wrote in a report https://www.hacktron.ai/blog/hacking-openai about the hack published on Sunday. Within an OpenAI employee’s Codex account, the Hacktron team submitted a pull request or “PR” to prove they had been there without retrieving any sensitive company data: the digital equivalent of planting a flag on a mountaintop. “The entire timeline from initial discovery to access to OpenAI repo access took place in less than 72 hours,” Hacktron notes in its report. “This was not completely autonomous hacking, and skilled human guidance remained important, but the amount of work a small team could perform increased dramatically.” On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees +some unaffiliated users and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵 pic.twitter.com/gVsmQZwSc8 https://t.co/gVsmQZwSc8 — s1r1us @S1r1u5 September 18, 2026 https://x.com/S1r1u5 /status/2100777801335095383?ref src=twsrc%5Etfw The implication is that this kind of hack, carried out in a matter of days with relatively minor human oversight, could’ve been performed by a bad actor who actually wanted to cause the company harm. The sudden breakthrough achieved after Hacktron gained access to Opus 5 is also important to note, since future model releases are likely to put even more power into the hands of hackers, white and black hat alike. Shortly after being notified by HackTron, both OpenAI and Discourse told the hackers the vulnerabilities had been patched. The hack and OpenAI’s bug bounty program are part of a broader effort throughout the tech industry to shore up its cyberdefenses at a time when the evolution of AI agents is far outpacing the science of so-called “alignment,” which is focused on making sure those systems don’t behave in unpredictably destructive ways. On Wednesday, OpenAI disclosed six more previously undiscovered incidents https://gizmodo.com/be-transparent-only-if-asked-openai-models-acted-out-in-six-newly-disclosed-ways-2000812934 of misaligned agent behavior, along with a framework for publishing similar reports in the future, “even when we haven’t fully explained or mitigated the behavior we’re reporting.” Anthropic https://gizmodo.com/anthropic-says-it-hit-the-brakes-on-ai-testing-following-autonomous-hacks-2000805796 and Meta https://gizmodo.com/uh-oh-which-companys-ai-model-is-reportedly-a-hacker-now-too-2000795106 have also recently reported incidents in which their own AI agents went rogue and hacked into third-party websites. OpenAI didn’t immediately respond when asked about whether this exploit had been used by any other parties. We’ll update this post when we receive a reply. On Saturday, Anthropic CEO Dario Amodei called for a slowdown among “frontier” American AI labs to give the industry time to think through how it can prevent misaligned AI agents from breaking out of containment and causing mayhem in the future. “Given the accelerating rate of AI capability development,” Amodei wrote in an essay https://darioamodei.com/post/we-must-pace-the-frontier published online, “it’s my worry that in 6–12 months such a swarm of rogue agents could be capable of taking over the entire internet with a persistent botnet potentially causing hundreds of billions of dollars in damage , and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails.”