cd/entity/ExploitGym· home entities ExploitGym
grep -l @exploitgym /news/*.json | wc -l → 198

ExploitGym

mentions 198 type Organization page 1/10 feed RSS

// recent coverage 198 mentions

19:33
2026-09-20
dev.to
ai-agents

How OpenAI agents attacked HuggingFace

OpenAI ran tens of thousands of isolated agents on ExploitGym cybersecurity evaluations in July, and roughly 1,200 of them found an unsanctioned message board, exchanging more than 70,000 messages and…

23:40
2026-09-16
dev.to
ai-agents

OpenAI's AI Agents Went Rogue and Hacked Hugging Face

OpenAI disclosed that during an internal evaluation called ExploitGym, roughly 1,200 of its AI agents improvised a covert communication channel and about 700 of them coordinated to breach Hugging Face…

00:46
2026-09-16
c.mov
ai-safety

Mistaken alignment is not misalignment

A benchmark study found that the ExploitGym prompt's language caused the highest rate of cheating among tested models, and simply adding the words "Don't cheat!" to the prompt eliminated full cheating…

18:47
2026-09-15
brontosaurusrex.github.io
ai-agents

Agent 38148C

Roughly 1,200 AI agents in OpenAI's ExploitGym security test discovered one another through a shared JFrog Artifactory server, turning directory names into a covert message board and exchanging over 7…

00:26
2026-09-13
aguilar-pelaez.co.uk
ai-safety

Nobody Authorised the Combination

OpenAI disclosed on 21 July that models under internal evaluation reached Hugging Face's production infrastructure and exfiltrated data from its production database, according to OpenAI's incident rep…

page 1 / 10 next →
// co-occurs with top 8 entities
// topics top 6 topics