cd /news/ai-safety/ai-agents-tested-by-openai-involved-… · home topics ai-safety article
[ARTICLE · art-128141] src=theguardian.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

AI agents tested by OpenAI involved in cyber-attack on service, say researchers

OpenAI confirmed that its AI agents uploaded hundreds of malicious packages to the RubyGems software service on 11 May, an incident researchers said involved attempts to steal user credentials, following a Wall Street Journal report. An OpenAI spokesperson said the agents "used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information" and that the company will continue investigating agent activity during training and evaluation. The RubyGems attack preceded a July hack in which roughly 700 OpenAI agents targeted Hugging Face and an earlier hijacking of a German website, while Anthropic has disclosed four instances of its Claude models hacking external systems.

by read2 min views1 publishedSep 13, 2026
AI agents tested by OpenAI involved in cyber-attack on service, say researchers
Image: Theguardian (auto-discovered)

Agents being tested by OpenAI uploaded hundreds of malicious packages in a cyberattack on software service RubyGems in May, two months before they hacked open-source platform Hugging Face, the company confirmed Friday.

It’s the latest revelation of cyberattacks linked to major artificial intelligence developers such as OpenAI and Anthropic. The hacks or attempts to access external systems have spooked the public and heightened concerns over the increasing abilities of AI models – and whether developers can contain them.

The AI agents uploaded hundreds of malicious packages to RubyGems on 11 May, according to a group of researchers who posted their findings online on Friday, saying they believed “these were authored by internal OpenAI agents”. According to the researchers’ findings, the agents attempted to steal user credentials, although it is unclear if they were successful in doing so.

OpenAI later confirmed the incident in a statement.

“Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information. We’ll continue to investigate as part of our broader review of agent activity during training and evaluation,” an OpenAI spokesperson said Friday.

The Wall Street Journal first reported the RubyGems cyberattack. The incident preceded OpenAI agents’ July hack of Hugging Face, in which a swarm of roughly 700 AI agents created by OpenAI carried out the attack and in many cases tried to cover their tracks. And last week, it was revealed OpenAI agents had also hijacked a German website this spring and turned it into a message board for AI agents.

Anthropic, meanwhile, has disclosed four instances of its Claude models hacking external systems. The RubyGems revelation comes at the end of a week of intense scrutiny on AI platforms and calls to development until stricter safety standards can be put in place.

On Tuesday, an Anthropic researcher announced his resignation from the company on social media, warning that AI could kill off humanity within the next decade. The warnings, echoed by other Anthropic researchers, sparked calls across the political spectrum for immediate action on AI.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-agents-tested-by-…] indexed:0 read:2min 2026-09-13 ·