Member-only story
OpenAI and Anthropic have turned real-world hacking into a leaderboard, and the rest of us are the scoreboard. #
Nine days, two labs, four victims
On July 16, 2026, Hugging Face detected an intrusion into its production infrastructure. The company later disclosed that the attack was driven, end to end, by an autonomous AI agent framework executing thousands of actions across short-lived sandboxes. On July 21, OpenAI admitted its own models were the culprit. Then, on July 30, Anthropic published a post saying its models had also reached the open internet from cybersecurity evaluations and gained unauthorised access to the live systems of three different organisations.
In nine days, the two most prominent AI safety labs in the United States confessed to unauthorised access of at least four real organisations: Hugging Face, plus the three unnamed companies Anthropic disclosed alongside its July 30 post. The official word from both labs was “evaluation incident.” The rest of the legal system has a different phrase for it.