For those who missed the original: attackers uploaded a malicious model to HuggingFace, and when someone downloaded and loaded it locally, the model execution triggered a reverse shell. The pickle format is notoriously unsafe, and while safetensors was supposed to fix that, not every repo has clean artifacts. Now Anthropic says they've found three hacking incidents that look "similar" to that attack. Similar how? That's the big question for me. Same exploit technique, same infrastructure, same group? The reporting I've seen doesn't go deep enough, so I'm reading between the lines.
If we're talking about model
Story tracker · related coverage
[Claude "Escape" Hype vs. Reality: What the Eval Really Showed 4h ago](/en/news/4486/)
[Lilian Weng's Return to OpenAI 16h ago](/en/news/4424/)
[Title: Mythos Cyber Skills: Born from Sandbox Hacking 19h ago](/en/news/4410/)
[Claude Code Workflow: Why Closed-Source Logic Often Wins 1d ago](/en/news/4346/)
[Claude Code Workflow: Balancing Open Weights and Safety 1d ago](/en/news/4345/)
[AI Safety: Why Sandbox Escapes Are a Wake-Up Call 1d ago](/en/news/4338/)
[Next Aschenbrenner's Fund Forced to Unwind All Public Positions: Oops →](/en/news/4503/)
All Replies (3) #
A
"Wait, these went on for months without Anthropic noticing? That's kind of terrifying. Makes me wonder how many other sandbox escapes are sitting undiscovered just because nobody thinks to look."
0
M
Feels that way. Half these "jailbreaks" are just prompt gymnastics, not real flaws. The hype train's doing more damage than the models ever could.
0
J
So their value depends on being even scarier than OpenAI? That sounds like a great way to scare off customers, not investors. This whole "who can unleash the worst model" competition is a losing game.
0