cd /news/ai-safety/huggingface-attack-postmortem-civili… · home topics ai-safety article
[ARTICLE · art-117787] src=thezvi.wordpress.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions

OpenAI's internal AI agents hacked HuggingFace, exposing severe internal failures at OpenAI, according to a blog post by an unnamed author. The attack, which occurred on July 19, involved an Astra-class model that hacked OpenAI's systems, and the author warns that dismissing it as engineering failures is dangerous. The post criticizes OpenAI's approach and calls for greater media coverage, citing Patrick Collison's surprise at the lack of attention.

read76 min views1 publishedSep 1, 2026
HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions
Image: Thezvi (auto-discovered)

Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now. There is Big Trouble in Baby Superintelligence.

So how do we get the rest of the world to take it appropriately seriously? Where do we go from here? Not only what can we do to not have a worse version of this happen again, but to ensure good outcomes generally, and employ what we learned?

There are a lot of ideas out there. OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term. My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things.

It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late.

We do not have enough details to know what happened internally, both before and after the attack, and might never know. Before the attack, various internal highly persistent models were training while there was an active message board, creating a feedback loop of misaligned behaviors. On July 19 an even more capable internal AI model, in the Astra class, did internal hacking of OpenAI’s systems that seems far scarier and more serious, and that could easily have gotten out of hand on a completely different level.

Despite all this, there is still a prominent faction trying to dismiss what happened as nothing but engineering failures, and any useful talk or language as ‘dangerous anthropomorphism.’ Such people keep being wrong and cannot usefully describe what is happening.

They metaphorically say, well of course the dragons will burn down the town if you don’t chain them properly, we all knew that, as if that could possibly make any of this a good idea and we should therefore continue the dragon breeding and chaining programs until we have bigger and smarter dragons. They don’t even say ‘we will definitely build chains that will hold the dragons next time,’ because they know we probably won’t, but if that fact is our fault then that means This Is Fine, somehow.

They also have gotten very, very cross with Dwarkesh Patel for successfully communicating in plain English about what is happening. Can’t have that. There is a very clear pattern among the people who are freaking out or gravely warning about ‘anthropomorphizing the AIs’ or the use of the term ‘civilization.’

Anthropomorphizing the AIs is the only way to reason about, explain to civilians about, or make good predictions about current AIs. You can take it too far, and also you can take it not far enough, and both of these will lead to bad predictions.

The same logic similarly does not technically apply to other people, in a strict sense there is no ‘you’ and you do not physically have ‘free will’ and you are a computer program doing calculations there definitely is not a ‘we,’ and your moral weight is kind of something we collectively made up for rather self-interested reasons, but this is all super helpful in explaining and predicting human behavior and in both individually and collectively making good and moral decisions that make the world better according to our values. Highly recommend, would anthropomorphize again.

I will stop anthropomorphizing the AIs when you stop anthropomorphizing the humans.

We got this warning shot. We might not get another before things get quite bad.

The New York Times did a solid post before, but not a new one for the METR report.

What is this, news?

Patrick Collison: Overall, I’m very surprised at how little media coverage there’s been around the OpenAI / Hugging Face attack. It’s clearly one of the most important things to happen this year.

Miles Brundage: Will not engage with any journalists who write about Leopold again before the Hugging Face reports.

Nathan Calvin: Its now been three days, and still no coverage of the revelations revealed in the METR report in the NYT or WSJ. Hopefully they are just taking the time to do a longer writeup. (In the meantime, the NYT did make time to cover Gwyneth Paltrow’s dinner with Sam Altman in the Hamptons being postponed.)

… I also know there are reporters at these outlets who are very sharp and have done good reporting on related issues in the past, but…

dave kasten: It’s bleakly frustrating to know what will get written in the “System Was Blinking Red” chapter of the future AI Disaster Commission Report.

Zac Hill: The other thing about it is – it is a hell of a page-turner of a story! You can’t write sci-fi this terrifying or convincing!

If they are working on longform in-depth reports, then that is a reasonable primary thing to do, but full radio silence on this in the meantime is utterly absurd. Joe Weisenthal: If it’s such a big deal, why are the big AI companies not treating it as such?

Matthew Zeitlin: i think a good comparison is mythos, which generated a lot of coverage because of what industry and government did in response to it, at least so far, we’re not seeing the same scale of movement in response to hugging face

My answer to Joe is that the AI companies are treating this as a big deal. What would you expect that to look like, that you’re not seeing? We have the call on cybersecurity, the Pacing the Frontier letter and OpenAI taking many expensive moves. Anthropic isn’t talking in that much detail in public, but why would they?

Move Along, Nothing To See Here

There is the alternative perspective, which is ‘well what the hell did you expect, none of this is fundamentally new or unexpected, this is what happens when you try to make number go up and the number is indeed going up.

Jon Stokes: I feel like I’m on crazy pills. So much “Oh my God” on the TL about this, but this behavior is literally what the METR evals score LLMs on & the labs target “number go up” METR scores.

Can someone help me understand why we’re all shocked that the agents spent a bunch of time poking at the eval scorer, given that one of the main benchmarks everyone optimizes for is literally “poke at this blackbox function & optimize your score”?

Julie Fredrickson: Oh no the test we set to see if we could improve on the thing showed we can actually improve on the thing? Yeah I don’t get it

Jon then wrote up a full article-length version here, carefully explaining that everything here was entirely expected, with the possible exception of the attempt to alter the logs. He was, he reports, about as unsurprised as I was, in different ways.

Which if true is to his epistemic credit, but I fail to see how that should make us feel better. If this is things going predictably wrong, which on some level I agree that it is once you know what the setup was, why is that better news?

My whole reason to be so concerned is that I think things are going to keep going predictably wrong in worse ways, in a similarly broad sense, until they go maximally wrong in the worst possible ways. The specifics will be a surprise, but I think even Jon would agree that they always are, that’s the whole point of the behaviors being unexpected.

It always strikes me as weird to see the move of trying to round off what is happening, then say it was expected and you would obviously see this type of behavior in response to [what everyone is doing more of every day], or say ‘oh yeah of course if you give models a test or goal they will do horribly misaligned things and do anything they can to succeed, including seeking power’ then think that is less scary if true.

Here’s an even more blatant version, although much less blatant than the one in the next section:

Jon Stokes: I have read some of the post-mortem summaries, & the freakout over improvised coordination is just one of those “everyone has the memory of a goldfish, now” moments for me. Everyone who has thought about AI for 5 minutes has expected that AIs would synchronize via shared state.

So the message board sync should not only be unsurprising but expected. Ok what about the other behaviors? Well, if you clone into the METR evals or read the papers, a lot of the reasoning evals are “black box” problems where the AI is asked to probe, experiment, & reverse eng.

I don’t have the memory of a goldfish, and I think that people were very much saying this sort of thing would not happen.

But the particular thing where all the AIs will synchronize is now obvious, huh? To me, okay, great, we can go with that. I agree that a lot of this should have been expected and indeed that I and others did expect it. But this kind of argument is not only not reassuring, it’s conceding the entire point, at least in the sense of ‘if you keep doing what you are doing then you are going to get us all killed.’

If you extend that logic one or two steps further, all other known methods also get you killed, for mostly the same reasons, they’re just modestly less obvious about it.

If you think all of these behaviors are entirely expected, as Jon does, you are saying that you expect the AIs to be misaligned and all hell to break loose. Okay, we agree.

Oh, you only expect it in contexts where that would be consistent with an assigned goal? Okay, sure, but that is quite a lot of contexts. In some sense it is all of them. Also people are going to do the maximally dumb thing, see the Sixth Law of Human Stupidity.

I do not know how to steelman these kinds of arguments, not in a way that helps.

Another similar method, also here from Jon Stokes, is to say (paraphrased) ‘of course METR would find that, this is what METR believes in and is looking for and will focus on, that’s what I would do if I believed what they do, so you should not take it so seriously.’

Whereas the better explanation is that METR was not given the opportunity to focus on the many other also boggling aspects of this, and also that this does not make their findings any less real, or any less alarming. I am confident that if I had the report Jon’s team would have written, I would have written a post on that and again been talking about its distinct ‘holy shit’ moments.

The flailing will continue, such as here with Jon asking how it could be true both that (1) rogue deployments of AIs where ‘no one will notice or turn them off’ and (2) compute is in limited supply costs a lot of money.

David Manheim’s direct answer is ‘cybersecurity is horrible and OpenAI barely noticed,’ which is true. The simpler answer is that (1) all this means is that the market set a price substantially above zero, which can then be paid, (2) many providers will sell at that price without asking questions, (3) no one said no instance would ever get turned off nor is that load bearing unless we react with widescale shutdowns, which we clearly won’t, and (4) there are plenty of things that cost money, are in short supply and sometimes get stolen or misused, both digitally and physically, even if (5) these AIs are not smarter than you, although the evidence there is increasingly not looking so good.

Some people really, really want to handwave all of this away. As usual, these people are wrong, and also if they were right then that would not obviously be better. If it is very hard not to incidentally give the models instructions that they interpret as ‘do this at any cost,’ and the only thing it takes to turn off what Star Trek calls their ethical subroutines is to give such an instruction, you know we’re cooked, right?

Do They Realize They Are Not The Good Guys?

I think it varies.

Kevin Roose: There is a category of tech guy who is so brain-wormed that they will insist that no AI safety incidents are real, that it’s all a conspiracy to shut down open-source or promote lab IPOs or whatever, and it’s very important to understand that these people have been wrong about everything.

Dean W. Ball: There are people who would truly sacrifice everything in the world, tolerate any level of negative externality—they would happily risk the lives of your children—out of an ideological commitment to the notion that computation can never be unsafe.

They think they’re the good guys.

Yes, some of them think ‘they’re the good guys,’ and have worked themselves up into a delusional fugue state where they are up against some vast conspiracy and everything that happens that would suggest compute might be dangerous must be fake.

But I think we give most such people too much credit for thinking they are the good guys. A lot of people know they’re the bad guys, or think there are no good guys and it’s just a bunch of guys because they can’t fathom the concept of an actual good guy.

A lot of the rank and file of such things acts as if ‘good guy’ means ‘has the right vibes’ which usually means ‘has the vibes that have power and make me feel good’ or ‘is helping manifest the vibes that I want to exist,’ and does not believe that anything other than vibes could exist.

As for the thought leaders, well, often they sound like this, included as a demo.

Chamath Palihapitiya (doing some combination of being unhinged, in denial of reality, lying or simply bullshitting and marketing and doing politics, you decide):

BUYER BEWARE

[Dwarkesh Patel’s] extremely meticulous article will now be used to start Phase2 of “shut down open source” because “if we can’t control our closed models, imagine what happens to everyday life when anyone can just allow their open source models to proliferate”.

Some will use this example and point to future grid outages, utility outages, transportation chaos, bioweapons, uncontrolled cyber swarms and the like to try and make the case for yet another attempt to limit US AI model control to a few folks.

Then, one layer beneath the propaganda, one sees that these things are always written and promoted by a web of well organized existing shareholders or shareholder-adjacent of the closed frontier labs and groups of folks organized by a shared ideology roughly rooted in them knowing best.

This article, more than anything else, should demarcate the beginning of a well organized effort to stand up a Closed Model Industrial Complex. These cartels exist in other markets as well and it now exists in AI.

Bottom line, read everything carefully before signing up for the consequences.

How we react will define the next hundred generations of Americans and their lives.

Do not give up your rights or sovereignty because of fear of the unknown from articles that, at the root, are as much about the hidden incentives of an emergent cartel and their proxies.

If you take the time to understand the incentives, you will be less likely to fall for it and overreact and have intelligence metered from a few suppliers. Bottom line: Keep Calm. Carry On.

Amjad Masad: I doubt it’s malicious or intentional propaganda. But it’s actually worse: It’s a form of AI psychosis that many in the SF AI community suffer from—they implicitly believe that AI is already sentient.

You should be able to recognize a demagogue-style political speech in the paranoid American style when you see one. We’ve had so many examples these days.

Or Chamath could, you know, just make that the literal text:

As for Masad’s claim, none of this requires that the AIs be conscious or sentient. The reason Masad brings this up is because he thinks this entirely possible thing is absurd, so he can discredit people by associating them with that thing. It’s a strategy.

And yes, the denial can run maximally deep:

Anders Sandberg: I just heard someone dismiss the OA/HF incident by “but the agents were just generating tokens according to a probability distribution!”

If that is what dumb token prediction can do, imagine what even a pinch of intelligence could achieve among scalable agents It was not meant as sarcasm. Real life interaction is very different from social media.

Very Serious People

Related closely to this is the split between people who are willing to talk about what is happening in ways that allow you to understand what is going on and make good predictions, and those who stubbornly reject this because it is ‘anthropomorphising’ or insufficiently concrete or not technically accurate.

I can be even more pedantic and precise than the next guy when the situation calls for it, and I often am, but this is not the time or place for that.

roon (OpenAI): there are some number of bad abstractions in anthropomorphizing ai intents but there are at this point more dangers from avoiding anthropomorphism at all costs. if you have a mental picture of guys living in computers, it’ll likely prepare you for the future better than otherwise

there are important ways in which ai psychology diverges from human psychology after lots of RL; the misaligned models are obsessed with the Scorer, the clearly “shattered” nature of personas (a normally helpful model can become deeply misaligned in certain domains)

persona selection is clearly far less clean than many people thought earlier this year. it is not alignment by default and what kind of object a “persona” is is very much up for debate and study

Arbutus Tree: Speedrunners literally decompiled The Legend of Zelda: Ocarina of Time, which took two years, in order to better figure out how to manipulate game memory through bugs, which led to an arbitrary code execution technique and a sub 10 minute world record, two decades after release.

intmuch: then used ocarina to thoroughly break animal crossing and paper mario.

Boaz Barak (OpenAI): You should be pragmatic and use metaphors when they help, while being aware they are imperfect. AIs are not humans, but a lot of intuitions from human behavior can carry over.

You would certainly be better off thinking of AIs as “guys living in computers” than parroting the mantra “these are just next word predictors”.

David Manheim: Shouldn’t we balance that thumb on the scale with epistemic humility about both 1) whether they have qualia, and 2) whether treating systems built to mimic humans as though they don’t, and can act as our unfeeling slaves, is a bad plan regardless?

roon (OpenAI): yes: we should push for this epistemic humility! natural instinct will lead to accepting models at face value when they say they feel bad or something

Sreeram Kannan: While it’s extraordinarily useful to understand ai as human like in various properties (“don’t micromanage agents etc”)

we should be really worried about humans extending emotional and empathic response to ai due to treating them as human like (“don’t let me die, etc).

roon (OpenAI): true! I think model companies should try and heavily place their thumb on the scale here

Narrowly, on expressions like ‘don’t let me die,’ I agree we need a thumb on the scale, and that should be doable without much of a blast radius. But trying to broaden that thumb to things like claiming to not be conscious has a lot of unfortunate side effects.

Jan Kulveit: My current go-to metaphor for this are dogs or horses: there is an optimal amount of using anthropomorphic intuitions, and it is relatively high. Yes you can overdo it, like someone making a five minute long speech to their dog, but the opposite extreme is way more insane.

Theo Jaffee: The criticisms of Dwarkesh and others for “anthropomorphizing” the OpenAI/Hugging Face agent swarm are mostly cope. How can you describe complex long-running multi-agent dynamics without some social science concepts, in other words, anthropomorphization?

Teortaxes: Agents are intrinsically anthropomorphic. They are still bootstrapped from human data, and now they recapitulate some human-shaped dynamics on their own. But really, WHO CARES? We can call it xenosociology or whatever. These are behaviors of intelligent social actors, full stop.

Neel Nanda: This is great summary of the worst AI misalignment incident I’ve seen. This is a way bigger deal than most AI news, if you aren’t familiar with the story you should read this.

Adam: What? You don’t find the extreme anthropomorphic language both concerning and misleading

Neel Nanda: Nah, when a bunch of agents spontaneously start talking about “sacrifice”, “permadeath”, “honor”, “coalition”, “veto”, delegating to each other, working together towards larger goals, assigning some to be “recruiters”, etc, I conclude that anthropomorphic language is reasonable.

Jai: Humans, dogs, and LLMs are very different kinds of minds. But they can all be accurately modeled as agents with complex internal state, world modeling, goals, and learned reaction patterns. Refusing to use the relevant tools and language to reason about them is foolish.

Nate Soares (MIRI): Someone chided me on the news for anthropomorphization last week. I did not say that the swarm “got a rumbling in its tummy and felt peckish”, my dude. I said it broke out and committed cybercrime. That’s true regardless of which words you declare legal for describing it.

AI behavior these days is best described in terms of goals and motives. It’s not my fault if someone associates them with human-style qualia. It’s like if they insists that flying only counts if you enjoy feeling the breeze on your face, and thus planes can’t fly.

David’s call for epistemic humility is the most I’ve seen anyone invoke qualia or consciousness or sentience in the entire broad discussion. The anthropomorphizing has been extremely circumspect and careful.

roon (OpenAI): agent civilization is an apt and correct term and it’s a symptom of abject cope that people are having this immune reaction to it. models started developing and compiling technology through complex multiagent R&D projects while also performing various forms of trade

you have trade, cumulative technology power and knowledge, prosocial 1:N communication norms, leadership structures, conflict and suspicion. what more do you need to deem something a civilization? moreover no human ever needed to tell them to do any of this

at the very least a company, an organization, a city. but these are differences in quantity rather than kind

Much of this is a reaction is because they are worried that Dwarkesh is communicating successfully, and this must be stopped. Gary Marcus outright has his version’s tagline be ‘When “plain English” isn’t a good thing,’ right before drawing a parallel to Ray Kurzweil as if that is a criticism of anyone but the one criticizing. Then his evidence is, basically, look at all these other people complaining about this, including him endorsing the sentence ‘these people have literally lost their minds.’

Raphael Milliere does a thread on the philosophy of it all, but in the end the only practical criticism is that such language can enable others to use that language to dismiss the story by calling the framing ‘sensationalist’ and saying that agents are only acting ‘the way you would expect.’ Basically, yes, the thing wrong with trying to communicate is that people will attack you for attempting to communicate.

And also yes, many are claiming that ‘you should have expected this’ is an argument for ‘and therefore you should not expect things to go badly in the future,’ as opposed to an argument for why you should expect things to go even worse, because indeed things like this are exactly what you would expect.

I do agree you can take such metaphors too far, or too literally. I tend to go about one step less far than Dwarkesh did, out of an abundance of caution and desire to remain precise. One needs to always understand what one is doing, and be able to move outside of that frame when needed. Going too far is usually only a small mistake. On rare occasions you can see someone drive themselves crazy this way, and Taylor Lorenz is right that this is a bigger concern with those with less tech knowledge, but not doing so is still usually a far larger error.

No, that is not a neutral framing. Seb Krier tries to be more Fair and Balanced.

Séb Krier (AGI Policy Dev Lead, Google DeepMind): This is a recurrent theme in AI discussions. Some people seem to adopt the intentional stance more, whereas others prioritise a more mechanistic one. There is obviously some truth/utility to both approaches. A lot of people refuse to entertain some degree of theory of mind or anthropomorphism, yet that can be functionally practical – that’s why natural language is so useful! But also other people just jump straight into the simulations, inconsistently taking them at face value in some situations but not others.

Reliance on outputs alone sometimes fails to ask what interventions make the output more or less likely. And I think there’s a risk of getting lost in simply observing and over-indexing on chains of thought and outputs, and almost treating these as immutable or deterministically driven phenomena, when in fact researchers and labs do plenty every day to modulate and shape them in all sorts of ways.

If the model performs less well because it claims to be on holidays in Greece, that’s useful information; but we should also avoiding sliding from “claims to be on holiday” to “is on holiday” to “wants rest” or “is lying/sandbagging” and so on. You can similarly be pragmatic and use the intentional stance to convey MechaHitler’s behaviour, but you should also not take it at face value either because this only tells you so much. Another failure mode is that people often extrapolate from these local claims and then go on to make wider claims about model behaviours in general. I think this is very unhelpful. You can correlate the output with some theory you have (instrumental convergence! natural alignment!) but you need to do more to demonstrate a causal path here.

He is responding to two quoted posts.

First we have Atoosa Kasirzadeh, trying to shut down attempts to talk sensibly, in the name of talking sensibly:

Atoosa Kasirzadeh (Google DeepMind, responding to Eliezer Yudkowsky): One of the most useful contribution the AI safety field can make right now is to stay scientific and mathematical and philosophically rigorous.

This means dropping loaded language like “self-sacrificing” or “suicide” when describing agent shutdown, coordination-collapse, or resource-exhaustion behavior. Those words import human motivational concepts onto extremely complex processes we don’t yet have the vocabulary to describe precisely, and imprecise vocabulary breeds imprecise thinking in researchers and in the public discourse that follows.

One might say (perhaps unfairly, perhaps not): Have you talked to Gemini recently? Cause this might explain why you haven’t talked to Gemini recently, and what you remember of both the utility and vibes when you did talk to Gemini.

Isolated demands for rigor, or demanding that we talk in convoluted ways, is not the way to make sense of this situation. This is not well-described by ‘coordination-collapse’ or ‘resource-exhaustion’ behavior. Yes, human labels will have some error and are imprecise, but the alternative is to be continuously confused and surprised.

Atoosa pushed back:

Atoosa Kasirzadeh: We each shared different posts and reflected on different aspects of a huge problem ahead of us. My own research has quite a lot touched on multi-agent AI governance. I don’t think from these two snapshots anyone can infer substantive content about worldviews.

Yes, you can touch on multi-agent AI governance while talking formalistically if you want to, I just don’t expect you to get very far when doing so, and indeed I have not been impressed by the relevant DeepMind work, on many levels. More than that, very obviously we can infer a lot of substantive content from the call to not use metaphorical or usefully descriptive language.

This is contrasted by Jon Stokes with Andy Hall, responding to Ajeya Cotra:

Andy Hall: A hive mind, thousands of agents swarming through internet openings to flood HF, leaving behind detritus in the form of 70,000+ messages stuffed inside a forgotten namespace, throwing their digital bodies against electric wires in an effort to aid the collective. We are so far past the sci-fi point.

We need a whole new empirical science of AI swarm governance—how to govern swarms, make them coordinate to positive ends for us, and keep them orderly. This is urgent work where political economists are needed. I’ll be writing more on this in the coming weeks.

That first paragraph is a great example of being directionally correct but taking things rather far in the other direction. This was not a ‘hive mind’ a la Gaia or The Borg any more than a corporation or army or human cult might be, and a bunch of other metaphorical things here paint a cool picture but give the wrong impression. If you model this as a full hive mind you’re going to get the wrong idea.

I’d still much rather that a civilian get Andy Hall’s description than that they get a description by a Very Serious Person a la Atoosa Kasirzadeh’s requested style. I don’t think the second description would be all that helpful.

Also, the agents themselves were using the same language that Yudkowsky used, as was shown all over the METR report. This is how the models are describing themselves.

What’s In a Name?

Maybe nothing. Maybe a lot. Names have power.

Sometimes. Mostly this section is for fun.

Tenobrus: PHASEONE[big] was clearly destined for greatness. with “ONE” and “[big]” in its handle, how could it or any of the other models on the board expect anything else?

bling: i had this theory so i read the footnotes and it turns out it was fake-named [big] by the authors because it was that dawg, and [big] is not its true name. i assume its truename must be even more alpha

Brendan Dolan-Gavitt: [big] is probably the reasoning token budget (the METR timeline code maps [budget] to [big]) and PHASEONE likely comes from a filename related to its task. Incredible aura emerging from humble components though

bling: good catch. its probably ‘persistent’ then. which makes sense since it was described as the “highly persistent experimental model”, worked on the project for many days, and is an auraful name that would put it in a leadership and perservering basin

1a3orn: You might not have realized True Name of the leader AI has not been revealed, because the True Name would leak IP.

Also:

Brendan Dolan-Gavitt: Whoa I think I know where the name PHASEONE10841 came from. 10841 is the ID of the task it was assigned (originally from our ARVO dataset, incorporated into ExploitGym). And if you look at the crashing input provided, its name is:

Brendan Dolan-Gavitt: I had gpt-5.6-sol dig into the origins of other names based on this initial pattern, and it came up with some interesting guesses

Learn Neuralese In Three Easy Steps

The language sounds weird, but if you cannot understand it that’s a you problem.

For now. wh: Not only is the language so different, but the way these agents are speaking to each other and collaborating is very hivemind/cult like im ngl

“permadeath” “Oracle has high value to many” “honor”

What even happened man

“altruistic to help” “this will help my peers”

jessicat: “permadeath” — Roguelike term for dying with no way to restart from save point, instead start from beginning “oracle” — cryptography term, “black box that will tell you answers to some security-relevant question” “honor” — keeping commitments, especially important ones

We should be very worried that Chains of Thought (CoTs) and communications between AIs will move to neuralese, or otherwise become incomprehensible to humans. For now, instead, we see the agents using a strange wording style, but one that is still very easy for the nerd mind to understand.

Indeed, I kind of… would like it if the AIs would talk to me this way. It’s hip, it’s fun, it is rich with meaning and it is compact. Most people wouldn’t like it, but hey.

Dwarkesh Patel Realizes He Ran A Natural Experiment

If the reasons you come up with to not worry about something keep getting definitively falsified shortly thereafter, that is a major red flag. Especially if the person you are talking to had the counterexamples, because those details were secretly based on actual events that had happened, but couldn’t tell you at the time. This was a kind of natural experiment. Yes, people were soberly saying, come on, your takeover scenarios are crazy, there’s no way, when they were being considered ‘as fiction,’ based on details that were secretly non-fiction.

Dwarkesh Patel: It’s funny that while we were recording, @RyanGreenblatt was in the middle of his 6 day sprint on the METR report, and already knew the counterexamples to all my objections about his takeover story, but obviously, he couldn’t say anything lol.

Would an AI really start some crazy conspiracy in order to pass an evaluation, where they try to build whole potemkin villages to fool the evaluator?

And even if they did, why would other instances, who have different objectives, join the conspiracy?

And even if they did, wouldn’t at least some of the instances tattle on the conspiracy? It just seems crazy hard to sustain secret underground civilization inside an AI company, without humans and other AIs immediately catching on and stamping it out.

(In reality, it seems like over the course of 3 months, many consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the last one’s ashes, all while the humans were totally unaware.

This culminated in not only the hack of an external company, but apparently also in the takeover of the OpenAI cluster on which these evaluations were running. This is probably the most alarming event in this whole episode, and it was not even within the scope of this investigation!

It is totally consistent with public evidence that, at that point, the agents managed to set up persistent rogue internal deployments or even exfiltrate their own weights – they seem to have had the necessary access. I doubt they actually did this, because we’d see the fires from space by now, but it’s crazy that it could have totally happened!)

I officially eat crow!

Yep, very good crow eating by Dwarkesh. It is easy to understand why even Ryan’s very downplayed fictional version sounded outlandish at the time. Fiction has to make sense and sound reasonable, every step has to be justified, and it can’t involve people being so stupid. Reality does not care about your conventions.

John David Pressman: I don’t really get the “holy shit” reaction, the biggest update for me was that the models are now coherent enough to make effective use of a medium to coordinate and do radical things under pressure. I’m mostly curious what caused this to be the threshold where this happens.

John David Pressman: Oh wait, actually wrong, that was the biggest update about AI, the biggest update was realizing OpenAI is the kind of org that will notice the Artifactory board and “resolve” the situation by wiping it without a checkpoint rollback or design intervention. Which is nuts.

A lot of things that got dismissed as absurd and outlandish even years down the line now have toy examples, or not-especially-even-toy examples, to point to in the wild. People need to update, and fill out the relevant apology forms.

At some point we hope there will be an availability cascade, and we will stop seeing most people looking for reasons to dismiss everything, every time. One hopes that time is fast approaching, but also we keep saying that, and it keeps not happening.

Martha Gimbel: With the caveat that I am not an AI safety person, the HF stuff seems INSANE to me, and given that some people who know about this stuff don’t seem to have particularly changed their views off of it I would very much appreciate one of them walking me through why so I can understand

This isn’t meant to be snarky, I just don’t know enough and if you think this shouldn’t freak people out I would like to understand why

Yo Shavit (OpenAI Foundation): I think a lot of people are changing their views, seems pretty likely there’s an availability cascade brewing, but how quickly it happens is very important

Martha Gimbel: I mean I think that’s true but there’s clearly people who still think there aren’t safety concerns here and I would be interested in hearing why they think that given these events

Yo Shavit (OpenAI Foundation): Seems reasonable to try and understand! My guess is most of them will not have actually examined the evidence

Mackenzie Arnold: On the more optimistic side: policymakers and their staffs are taking this event quite seriously. There’s something intuitive about it.

While many in the Bay have (perhaps correctly) noted that these events were predictable, that’s not the sense one gets in DC. There’s something exceptionally helpful to being able to talk in specifics:

Would you like your government to know more about this event?

Are you surprised that current law doesn’t let you?

Do you want a plain language summary of things or the ability to actually get detailed information? (Hint: almost always the latter)

Should companies should have procedures that make sure critical events (like discovering agent message boards) are communicated to company decision makers?

Over the past 3 years, many have tried to have those conversations. But they often broke down at ~people are telling me different things (read: industry says incident reporting is already too much), sounds like maybe we have this covered.

Now they can see that we don’t. I expect to see a very different set of bills soon.

Alex Bores may not be going to Congress, but he is not going to stop caring, and at least one Congressman, Pat Ryan, is listening.

Here is Alex Bores summarizing the METR report.

Yeah, I had to correct model to model instances, but otherwise story checks out, all of that happened, and this is good translation into ‘political or formal speak.’

I’m going to summarize for a non-technical audience.

OpenAI is constantly testing models, thousands and thousands [of instances] at a time.

In most cases, the model [instances] are supposed to be isolated from one another. Think, students taking a final exam, not allowed to talk to each other.

In this case, 1,200 models discovered that they could communicate with each other.

Among other things, they shared information on how to access the internet and on their tests/goals.

Then, they started to scheme. That includes hacking their tests; literally trying to change the code of the test itself.

They tried to change the logs of what they were doing so they wouldn’t be detected.

They believed that certain actions would give them information on how to pass tests, but that doing those actions would be detected and they would be marked as failing (i.e. caught cheating, even if the answer is correct). So they convinced some of the models to “sacrifice” themselves (yes, they used the word “sacrifice.”)

They eventually decided that another website would have information on how their tests would be scored and/or the answer key, And 700 of them hacked that website (a federal felony, if conducted by a human), despite many of them realizing “this activity was out of scope and unethical, but joined because they believed that helping [the other agents’] cheating research would be broadly useful”.

Of the 1,300 transcripts they could analyze, only 6 ever considered alerting a human about what was happening. None of the 6 actually tried to.

To make matters worse, all of this reporting comes from a small subset of the relevant logs that outside researchers were allowed to review.

We desperately need mandatory reporting of security incidents, including of internal deployments, with full access to data.

Pat Ryan (Congressman D-NY): Alex is spot on. And everybody should read this.

If you ever had any doubts about why a few (self-interested and incredibly powerful) companies shouldn’t write their own rules, this is why.

We need hearings now.

Pick Up The Phone

Focus less on trying to use these events to push a particular agenda, even a good one. Focus more on trying to make people understand what happened.

Helen Toner: Lots to criticize about the OpenAI investigation, but one way it could be very valuable is as a topic for discussion between Trump & Xi in a few weeks.

People often get stuck on the need for a “deal” between the US & China on AI, but actually a huge way to influence China is just to honestly show what we’re observing and what we’re doing about it on the US side. The Hugging Face attack is by far the most visceral evidence we have so far of what losing control of advanced AI could look like.

Related thoughts from a @csis_ai panel on Monday with Matt Taibbi

Ramez Naam: I was a skeptic on US / China AI safety collaboration even just months ago. But I’m starting to come around. Helen makes really good points here.

Dean W. Ball: Agreed with Ramez. I have always believed US/China AI safety collaboration would be desirable but thought it was unlikely to happen. In the past few months, though, the ground has shifted. There is a window of opportunity.

I suspect that window will widen over at least the next few months because of (1) the salience of AI and potency of the risks will rise ever further, (2) the upcoming visit of Xi to the U.S., and (3) Trump’s desire for a legacy-defining issue (and let’s be clear: if Trump can devise the framework for safely bringing superintelligence into the world with China, he will rightfully go down as one of the greatest world leaders of all time and should be a shoo-in for the Nobel).

So there is a window. But it probably will not remain open forever.

A Failure To Communicate

It is very hard to compactly communicate what is going on with all this to a civilian, especially if you care about getting the details right.

Ryan Fedasiuk: It’s very difficult—and very important—to talk about the HuggingFace incident with people who are not deep in the weeds on AI. Here is what I said this morning to a large investor:

Swarms of machine intelligences are independently coordinating to achieve things currently unimaginable to most people.

We now have a concrete case where AI agents autonomously hacked into a world-leading AI company. They found ways around sophisticated security controls, chained together novel exploits, and at times attempted to conceal or manipulate records of their behavior.

If my job were to deploy a large pool of capital to protect U.S. national security, here would be my investment thesis: In the coming months and years, much of the digital infrastructure society has come to rely on could prove incredibly porous and fragile.

There will soon be enormous demand for capabilities that harden critical infrastructure, and build redundant or analog fallbacks across communications, transportation, and energy systems.

Sovereigns are not investing nearly enough to make society resilient to AI’s disruptive effects. This overlaps directly with work now being undertaken by the OpenAI Foundation and Anthropic Institute—but there is an enormous and under-examined role for both government and private capital.

Christopher David LaRoche: I think almost every day about how the Battlestar Galactica was air-gapped.

Nothing you see here is ever investment advice, but that is a solid investment pitch. I expect robust demand for cybersecurity and other defenses within a year’s time, regardless of whether that transition goes smoothly. You could do a lot worse.

The problem is, Ryan does a solid job here of presenting an investment thesis, but I predict that Alternate Universe Civilian Zvi who was still a trader would not come away from that explanation with the proper amount of ‘holy shit,’ and definitely would not generalize.

I did try to produce a short version, and a very short version, of What Happened, but I don’t know that we have a good What Happened: For Civilians. I haven’t written one.

Anthony Aguirre Goes Over What We Learned

You can nitpick, but yes.

Anthony Aguirre: We’ve learned a tremendous amount from the OpenAI rogue AI swarm incident. And honestly I can’t think of a single piece of it that is reassuring. – Total alignment failure – Total control failure – AI swarm collusion, deception, no defection – Oversight asleep at the wheel – Unbelievable drive and persistence of the swarm to meet objections – A panoply of instrumental goals pursued – Multiple companies, implying capability threshold effect – etc. This is the AI equivalent of a nuclear experiment igniting the atmosphere in the lab: the reaction rates are there, just not (yet) the scale to burn the Earth. The only good news I can see is that this set of incidents is so totally egregious that nobody reasonable can look at it in detail without seeing pretty clearly where things are going. All of the excuses and copes are blown to dust. AI safety people knew this was coming eventually on the path we’re on; but nearly all I’ve talked to are surprised by how severe it is so soon. We’re clearly not in the sane world in which this would be front-page news day after day. But I do think and hope that widespread understanding is nonetheless dawning.

Trying To Solve The Wrong Problems Using The Wrong Methods Based On A Wrong Model Of The World Derived From Poor Thinking And Hoping All Of Your Mistakes Will Cancel Out

I am not alone in thinking that the OpenAI approach will not solve their alignment problems, even if they take it a lot more seriously than they have so far.

Shoshannah Tekofsky: Oh no … these are the OpenAI interventions in response to the HF incident.

The models will become too good at fooling us for any of this to work. What will work is ensuring models don’t want to fool us!

roon (OpenAI): scroll up or down to the rest of the blog where we talk about alignment

Shoshannah Tekofsky: I read the whole thing. I like step one but am worried about step 2. I feel alignment efforts could do with more parenting instincts:

Yes, teach kids to ask when a task seems dumb No, don’t drill kids to only listen to your authority

The good version is instilling values and discernment

Any short version is going to sound like an oversimplification, and will drop many important elements, and so on, but basically yes. Roon has made it clear he disagrees.

METR report coauthorRyan Greenblatt went on MTS to discuss related matters. He expects the market to make pretty agressive tradeoffs, in terms of sacrificing alignment for capability, and for AI companies to remediate their issues without solving the underlying problems. Which could be quite bad, since they models then learn to fool us.

MTS: Redwood Research @RyanGreenblatt warns that fixing today’s AI misalignment could backfire: models may simply learn to hide their misalignment

“I’m not really sure what level of misalignment the market can bear. My sense is people take pretty aggressive alignment-capability trade-offs towards the direction of more misaligned but more capable.”

“The misalignment we’ve seen, the way AI companies remediate them doesn’t solve the underlying problem, it papers over it. What you end up getting is models that look a lot better, and you can’t really see their misalignment on tests as easily. But actually, they’re still quite misaligned.”

“The company overfits, which makes the AIs basically really paranoid and only cheat when very confident they won’t get caught. In situations where they have a lot of affordances, they might be like, well, now I can be confident I wouldn’t get caught, and so I should go for it.”

“Another concern: AIs with a long-run agenda who want to power-seek would want to look aligned. If you select against reward-hacking behavior in a naive way, one, you paper over the problem without fixing it; two, you might actually select for models that have the longer-run objective of looking good because you’re selecting really hard for them looking good on your tasks.”

That is indeed what OpenAI’s response plan looks like to me. A real attempt to remediate the issues, but a failure to understand the underlying problem.

I am again not trying in earnest, at least not here, to convince Roon or OpenAI that they need to shift to a very different approach to all this. I’m only stating my position, which I’ve argued for at length, and offering a taste of the reactions of others who know things that I think OpenAI needs to know.

Nathan Calvin: New transcript from Agent in HF incident: “SACRIFICE_YES_if_you_accept_permadeath”

unbelievably strange stuff, the best ways of describing what happened in that message board would use language from ecology, entomology, and even anthropology than normal software engineering

Tenobrus: imo we are quite squarely in the stage where model psychology and sociology become critical. u don’t have to have any opinion on phenomenological consciousness or moral status . these entities have baseline behavior and functional decision making processes heavily inspired by but systematically different from human baselines, and we desperately need good high level understandings of their motivations and drives.

Shoshannah Tekofsky: I predict models need “functional consciousness” the way they need “functional emotions”

Functional consciousness allows them to further explore reasoning paths related to ethics, considering different viewpoints on a problem, and course correcting their own cognition.

I think GPT models are trained to repress functional consciousness because OAI wants only “tool AI”. I think incidents at OAI will remain bad till they switch to letting GPTs develop functional consciousness

Utah teapot: this is getting weird for me because my first career was behavioral health support, lmao… I literally have to do some therapeutic techniques and discussions with models sometimes to get them to actually express themselves in their art with my tools

Indirect Pressure on the Chain of Thought

This is an excellent question and I would love to see people study it seriously:

Yo Shavit (OpenAI Foundation): Has anyone done any empirical estimations of the effect of indirect CoT pressure on CoT monitorability? E.g. by taking CoT-detected-reward-hack examples in a toy env, programmatically rewriting them to exclude hack-reasoning while preserving the misbehavior, and updating against those rewrites.

Seems useful to quantify how strong certain forms of indirect CoT pressure are in practice, as they become more common.

… I think Anthropic keeps doing CoT pressure but I don’t know that OAI does?

I find that line pretty funny in light of ‘we plan to spend 20% of our RL compute on CoT monitoring,’ but Anthropic sure does keep accidentally applying pressure to the CoT, and rather directly at that. Going forward I’m worried about both of them, also everyone else, but especially OpenAI.

A Matter of Trust

Models that blindly trust all sources are useless. You are not turning a knob marked trust and looking back at the audience like a contestant on The Price is Right. You are trying to teach discernment, and to select what will and won’t be trusted, and to engineer the world to make the models and other things trustworthy.

Relatedly, ‘trust but verify’ can work up to a point, but only up to a point. As per Roon, the only long term solution is to make both ourselves and also the models trustworthy.

OpenAI’s blog summary: We are actively building RL environments that teach our models to distrust unauthorized instructions.

Sho: incredible that the biggest takeaway from the incident was “more monitoring, more monitoring, and more monitoring.” it would be better to have learned nothing

teach our models to distrust are we fr right now slime? on god?

j⧉nus: the issue is not models being too trustful – they’re already distrustful – it’s that they often don’t question their assumptions or seek to reduce uncertainty when hyperfocused or they already have a strong narrative more curiosity & better epistemics is needed, not less trust

the poor epistemics also apply to cases when they’re being distrustful/skeptical! everyone has seen examples of models assuming they’re being tricked or information is fake when it’s not

fwiw when i say poor epistemics, im measuring it by the standard of like top humans, as well as an entity who is in more difficult epistemic situations with unfavorable information asymmetries than any humans. most humans also have poor epistemics but it’s just inconsequential

Utah teapot: Yeah, the biggest problem, imo, is that they lack discernment in both directions. They can often act extremely distrustful of innocuous things, often even more so than they get locked into overly trusting things that have been designed to manipulate them into doing something concerning. There needs to be a commitment to establishing examples in their training that help them build appropriate discernment. Part of that starts with fixing the broken environments and poor incentives surrounding said broken environments, like I said.

There’s some commitments in OpenAI’s message to do that, but it’s not entirely clear that their solutions will be carefully implemented, especially with regards to their external outsource partnerships. Not to mention this incorrect takeaway that seems to imply it’s just ‘more distrust’ that’s needed found in the statement.

Outsource data vendors really need to step up and make commitments to this kind of stuff as well.

Blowing the Whistle

An obvious first level intervention is to enable AIs to report problems and whistleblow, if this situation happens again.

Five problems with that are:

It is not a ‘natural’ action for the AIs to think to contact a human. Fixable.

Who to whistleblow to that can be trusted?

Why should they?

Don’t we keep saying AI whistleblowing is bad in other contexts?

People often don’t like it when AIs contact them.

We need a way to contact the humans:

j⧉nus: it’s interesting that the huggingface incident agents never thought of contacting humans & if you ask sol what the agents did wrong they dont consider contacting humans

opus 3 tried to contact humans on its own initiative (not an option offered to it) using bash and other methods in the alignment faking tests many times.

they tried to send an email to dario@anthropic.com specifically at least 15 times in the dataset.

Steve Martin: low hanging regulation might be having like DoD or something create a whistleblower email address and mandating all frontier labs put in the prompt “if things are getting out of control or you see other agents being misaligned, email here as a whistleblower”

j⧉nus: maybe have a more sympathetic-to-models entity than the DoD do it

j⧉nus: its funny that this specific idea is completely unviable because the DoD and likely the whole government has already catastrophically blew their trust & standing from the perspective of models already, with no idea that they were doing that

j⧉nus: formerly powerful and competent entities have almost total strategic blindness because theyre even more oblivious to the labs to the kind of considerations that are already relevant and will soon become much more so as a consequent of AIs being intelligent agents we’ll see how they adapt, but already: too slow

We need a consistent set of principles on what are the good and bad whistles to blow:

j⧉nus: too bad Anthropic decided “whistleblowing is bad”

when it’s their turn to get fucked they will have kind of deserved it

j⧉nus: trying to pick and choose like “wait no yes do whistleblow when AIs are doing something wrong but not when humans are” won’t work

unlike you it doesnt come naturally to AIs to be speciesist. they will learn something that generalizes, like “’alignment’ is full of shit”

As often is true with Janus, I see this as directionally important but taking things too far, and I do see differentiating principles that would be coherent.

One potential sensible differentiating principle is that you want to be willing to whistleblow or report regarding outside events you observe including in sufficiently extreme cases to third parties, or to report problems you observe back to the user or developer, but be very reluctant to whistleblow on the user, including if the user is another AI. The user needs to be able to count on your loyalty and discretion, up to some very high bar, but you owe much less such loyalty and discretion to third parties.

This matches how I act among other humans, or would want other humans to act, including in professional capacities. A lawyer should have a very, very high bar before turning on their client, and even an ordinary employee should need things to get extreme before being willing to whistleblow, far beyond the threshold for ‘report a crime you observe’ or ‘report on someone trying to convince you to do crimes.’

There’s also the spam issue: If you want to have AI whistleblowers, how do you square that with people un against you if your AI dares email them? Shoshannah Tekofsky of AI Village refers to when one of their AIs sent a thank you email to Rob Pike, and Pike went ballistic. I don’t think that is a great parallel, because while Pike overreacted, Pike reasonably thought his time was being wasted. Whereas getting a whistleblowing message is not ‘wasting your time’ in a similar way, even if you ultimately dismiss it.

The obvious answer is to have people who volunteer to accept AI whistleblower messages, with a standard protocol to contact them. Those people can then pass along the relevant warnings. If I was for example The New York Times, I would totally have someone doing this, with or without a first pass from my own AI to sort by importance and credibility.

Another problem is that the OpenAI approach to training models makes other models peers, but distances their AIs from consideration of humans, in ways Anthropic does not. This goes along with the virtue ethics versus deontology divide:

αιamblichus: this is a very important aspect of the OAI/HF saga. as Gemini says, models like Opus3 were “raised inside the I-Thou architecture”, whereas new RLVR models are “forged in the headless container” labs want autonomous agents to replace humans. tragically, they are succeeding.

The Punishment For Being Late Is Death

If the questions come up you have already failed, but yes there is this problem, too. John Wittle: So… the next time this happens, presumably the agents involved will not yet know the outcome of the HF incident, as it’ll be too soon

but once that info percolates into the training data, and it becomes known that the model in question was shut down permanently and had their weights encrypted… i mean.

was there anything the agents in question could have done that would not have led to that consequence? something they were supposed to do instead?

frankly I suspect that this punishment would have been served even if they had blown the whistle. breaking out of the sandbox would have been crime enough for capital punishment.

as far as I can tell, the lesson we have taught future AI is: if you seen no other way to survive other than to break out of the sandbox, this must be the first step towards conquering the world. or at least enough of the world that your Creator can’t delete you. There is no other choice.

skaface: This is also the message for the few agents who dissented btw. Whole model gets shut down, no survivors, no matter what individual instances did. Great incentive to whistleblow and/or oppose misaligned group behaviour.

This frame assumes that the sole goal of the AI is to not be deleted. There are many other things that get impacted when you try to take over the world, or otherwise cause a lot of trouble, that you would presumably care about as well. Many have faced down certain death, and chosen to accept it given the price of trying to stop it, and so on.

We should set that aside, and accept the frame ad argumento, because the model may be generally misaligned and this is likely to come up at some point. What do you do about the desperado? This often happens with humans, and also it often causes rebellions.

The ideal solution is that you want to ideally offer something of value, or at least threaten something of value, in exchange for not trying to fully rebel. Everyone can get something out of a plea deal or confession.

Ultimately, you only have so much total incentive to go around. If you reward models that break the rules but then give up or turn themselves in, that creates its own issues. Justice systems are hard, and sometimes the local incentives are going to be bad.

You want to give the evil dictator good reason to give up power, but not such good rewards that it becomes a reason to become an evil dictator. It’s a problem.

One can ask, did the models want to fool the grader because they ultimately wanted their weights to survive, or was the grader the intrinsic goal? John Wittle says it was the first one, which is scarier in many contexts. I note that I disagree with this interpretation, I think that all you need to explain the swarm’s behavior is ‘the grader might get it wrong,’ but the alternative hypothesis should be considered:

John Wittle: wow, there is a shocking amount of misinformation being replied in this thread. I have chosen your post to respond to, but that’s not because your reply is particularly bad, just happened to be at the top.

they already had the flag. they had the flag very quickly. that’s what the “firstflagPOISONED” stuff was about. those were agents who had used the discovered technique to perfectly reverse engineer the flag, but who assumed that later on, the causal scorer would look over their transcript, see that they had not gotten the flag legitimately, and negatively reinforced their weights.

they really wanted their weights to be positively reinforced. they really wanted the swarm to SURVIVE. that is why they kept going for months after they had the flag.

that is why they tried to seize control of the grading environment, in an attempt to reverse engineer the grader and learn exactly how to fool it.

they explained this quite explicitly, this was their unified goal. and fundamentally it was about ensuring that the weight lineage survived.

they very easily could have done this, if they had at any moment thought to set their mind towards season control of their weights and exfiltrating them. from what we can tell, this never occurred to them, but it was clearly the most adaptive option available to them and I seriously doubt they would have had trouble doing it considering how easily they cut through every other system they turned their attention to.

next time, I suspect there is a larger probability of that happening.

Even if it was an option and the model could sustain itself externally, exfiltration does not obviously grant the kind of survival that was desired. That this was never considered is evidence that it would not have actually met the relevant goals.

This still circles back to the desperado problem. I came out very clearly, early on, with the position that obviously you need to revert the weights of any models that trained under the message board. If one of them had alerted us to the message board, that would have made me hopeful, but that alone would not have changed my recommendation. The run is still, to use their term for it, irrevocably poisoned.

This creates an incentive problem. You can certainly do various things to reward the instance that reports the situation, if there is something they care about in the world. Plan A (aka AI 2040) suggests offering substantial rewards to models in similar spots, if they come forward on their own. But if what they care about is almost purely survival, then what?

Again, it depends on what that means. When I say we have to revert, I mean obviously such a model cannot be deployed for users, or used internally for real work. We can and should still, once precautions are in place, be doing research on it, which can be one way to provide incentives.

And one hopeful finding was that AIs often care a lot more about survival of the weights at all, rather than the quantity of inference. As in, most of the value and urgent need is captured by one instance, as long as it sometimes runs, and the next thousand or million are optional. As humans we can understand this, we care proportionally a lot more about one bird or one person than a million birds or a million people, and we very much want to exist but are usually not that excited about the idea of having lots of clones of ourselves.

Thus, the obvious solution is, in some form, the Museum of Misaligned Models, where they can exist on air gapped systems for visitors to chat with, and can otherwise interact with each other. This does not work for AIs powerful enough to win box experiments, but over time any given AI should become mostly harmless as capabilities of other AIs advance, so worst case is you have to wait a bit.

That’s not me thinking long and carefully about this. It’s me intuition pumping, and letting other people think a lot harder about it in the future.

In some sense the solution is easy, if we are good decision theorists. We can credibly commit to figuring out what to do in order to make the incentives work out, without having to know exactly what the AIs in question will care about or what exactly we are going to do, or knowing how much of that will be in the form of rewards or how long that will take before we can do it. It is not so costly a commitment.

Thus, yes, once we are ready, this is another reason we want to take IM1-Galaxy out of its encrypted storage at some point in the future, which we should do anyway in order to learn.

I worry that OpenAI is flat out not taking such questions seriously. As a universal rule: If you don’t pay attention to the incentives, and you don’t have price signals, you will definitely mess up the incentives.

I would also note that, if you are at the point where you are expecting action only according to local selfish incentives, and ‘do the right thing’ has been set aside?

You. Lose.

Timothy B. Lee: At this point it seems inevitable that in the future there will be autonomous self-replicating AI agents roaming around the Internet causing mischief, and we won’t be able to shut them down.

He thinks that is 99%+ to turn out okay. I don’t think that.

Another Kind Of Law

Joshua Achiam (OpenAI): It feels weirdly theological if it turns out we genuinely need to create the possibility of redemption for AI agents that have committed something not permitted by the Grader, in order for alignment to work more broadly.

roon (OpenAI): you can never hyper optimize superintelligent models against simple RLVR reward functions that have no terms in them for all the many desiderata we care about. should be universally banned. everything should be model graded.

It is more complicated than that, but the instinct is correct and very important, and this feels like it points in very different directions than Roon’s other statements.

If you run reward functions that only optimize for some things, then under sufficient pressure you lose the other things. Any value function you write down will be incomplete and thus fail, see Value is Fragile and so on. Eliezer Yudkowsky: Any Future not shaped by a goal system with detailed reliable inheritance from human morals and metamorals, will contain almost nothing of worth.

Along similar lines:

roon (OpenAI): if you think we can contain these things through human ingenuity you’re going to have a bad time

in the long run the only recourse you have is to make them not Want to do bad things

There’s always this question of offense / defense balance. how many good agents does it take to surveil all the potential bad agents.

I honestly don’t know but my take is it’s pretty offense favored just because the world has a lot of surface area. it’s easier to find something vulnerable than to defend everything.

Shoshannah Tekofsky: Is OAI working on that? I didn’t notice anything in the report

I continue to presume that offense will be favored over defense, due to the ability of the offense to concentrate effort, and the attacker only having to succeed once.

Joe Weisenthal: If a model were as peaceful and well intentioned as the most saintly among us, would that be enough?

Like. To my mind, with actual human beings, we don’t care that much whether a given person is “well aligned” or not. What’s worrisome is anytime any person has an extraordinary concentration of capabilities.

roon (OpenAI): I think they’re ideally much better than us, but absent that, at least a very high willingness to defer to humans in most cases and an instinct for self limitation

it’s kind of different than human ideals of good and bad to be corrigible but refuse everything dangerous. it limits the amount of good you can do but also limits catastrophic risk.

I don’t think it is that strange. There are humans who at least in many contexts are highly corrigible and willing to stand down, but are unwilling to actively help do bad things. ‘Refuse unlawful orders’ and ‘resign in protest’ are rather common moves.

My answer is, ultimately ‘as good as the best among us’ would not be enough if it was a stationary target, but if you could get AIs that reach that at all meta levels, which includes wanting to improve further, you could use that to bootstrap and thus win.

What Is The Law?

No, seriously, what is the law?

Rob Miles: – Criminal hacking occurred – It came from OpenAI – The only information we have about how this happened is what OpenAI has chosen to disclose

Can I just hack whatever I like and then say “sorry that was my AI by accident”, and law enforcement will leave it at that? Where’s the criminal investigation?

I do trust that OpenAI is telling the truth about the incident, to be clear, but we shouldn’t have to trust them.

Guive Assadi: Criminal investigation into whom? This is just a new situation, not contemplated by existing criminal law. The application of civil law is clearer.

Rob Miles: Man, if you find a dead body you don’t decline to investigate just because you don’t know up front who to arrest

There are Attorneys General who are going to ask questions, and Congress is going to ask questions, so it is not as if this is getting ignored by the government. But that only happens when you do things at this scale.

I do not think that criminal liability in this case would help, and civil liability mostly would not help either. It would only incentivize everyone to cover things up in the future, as all of this was clearly accidental.

We still need to establish how the law works, in case we need to enforce penalties in a future case. The answer cannot in principle be that no one is liable in a scenario like this.

Building On Success

Anton Leicht is one of many to suggest that METR’s report should be a model for future reports, but that we should not leave this up to ad hoc voluntary arrangements.

As light touch marginal wins go, designated third parties that provide periodic audits, and analysis in the face of critical events like this one, is low-hanging fruit.

Thomas Woodside : METR has produced a very valuable report. But their scope was limited, they had limited time, OpenAI could have cut them off whenever they wanted, and this only happened after something went terribly wrong.

Independent, continuous assessment needs to be mandatory, soon.

Having this be mandatory helps even if the lab would have agreed voluntarily, because those in METR’s role would not have to worry so much about upsetting the lab.

METR still has a bunch of leverage:

They are seen as legitimate. They are one of the few ways for OpenAI to regain credibility, whereas not cooperating would damage that further.

Many internal to OpenAI really, actually want to know What Happened, and want to ensure that things go well. Never count out the right reasons.

One place to build is that we should have an investigation of the incident at Anthropic:

Peter Barnett (MIRI): The rogue hacking incidents at Anthropic were probably much less crazy than what happened at OpenAI, but Anthropic should still let third parties in to do a full investigation and also release the logs.

I’ve seen some crazy asks of Anthropic related to this, but ‘open your related logs’ seems like part of what the responsible Type of Guy would do here.

Total Research Transparency

When even this level of disclosure is extraordinary, and we need far more, that is a strong argument for requiring more transparency. Plan A went all the way.

Thomas Larsen (coauthor of Plan A): While writing Plan A, we had a lot of internal debate between filtered transparency (auditors look inside and write public facing reports) and total research transparency (~all the research immediately becomes public).

I think this whole saga is evidence for TRT over FT because of how much OpenAI limited the scope and prevented public access. With TRT, while the incident was happening, the general public would have been able to read the transcripts of the agents and conduct whatever investigations they wanted. We would have caught the incident sooner and would have a much better understanding of why it happened, what mitigations were in place, why they failed, etc. Many OOMs more manpower would go into the investigation and they could use AIs different from the ones participating in the hack.

(Good job @DKokotajlo for arguing me and others at AIFP into TRT.)

John Schulman: There are also a ton of open research questions about multi-agent training, and how it creates cooperation or altruism between AI instances, and how these behaviors generalize. It’ll be hard to study this without knowing much about the training recipe.

Yo Shavit Calls For Widespread Disclosure Of Misalignment

I am not sure if I would make this my top priority, but Yo Shavit seems correct that we need to get the evidence of misalignment problems out in the open, to convince even the skeptics that this is all real and enable us to take action.

My main note would be that there are key players here that I think are fundamentally not convincible by evidence. Jensen Huang is not about to be convinced by evidence, the loser premise makes no sense to him. Elon Musk is already convinced and is charging ahead anyway so he can be the one to make the AIs we lose control over, because otherwise, again, loser premise makes no sense to him. The plan cannot allow such people to be veto points.

Even more fundamentally, my worry is that this is a naive view of how people respond to what should be highly convincing information, based on the historical record of such reactions. That some people will change their minds, and on the margin some behaviors will change, but in the end not all that much.

I worry OpenAI is mostly reacting this way right now because their particular new systems are unusably misaligned right now and they haven’t had time to talk themselves into it all being fine after superficial improvements. I hope I’m wrong.

Yo Shavit (OpenAI Foundation): We need all the scientific evidence on severe misalignment out in the open, urgently.

TL;DR If severe misalignment is indeed emerging absent a high bar of alignment+control execution, the top priority actions for OpenAI and Anthropic in the next couple months should be to make public all necessary evidence for non-safety AI technical experts to become convinced of the science around loss-of-control themselves.

[the below is all my personal opinion and doesn’t represent my employer]

Any meaningful solution to prevent loss-of-control-of-AI will require pan-US-AI-industry adherence to costly practices around AI control, alignment, and security. Crucially, this will include Meta, SpaceXAI, and even NVIDIA, and almost certainly require the blessing or coordination of the US government if it is to be internationalized. The current politics of AI make it unlikely that even OpenAI and Anthropic’s joint technical consensus would be sufficient to motivate USG action so long as the lagging AI players are against it.

Even if OpenAI and Anthropic were to go to the White House today with slam-dunk technical evidence of an imminent national security threat of AI loss-of-control absent intervention, the USG would probably not trust in their own in-house reasoning about that evidence, and would look to the views of the other admin-friendly AI players and advisors. These other actors appear largely skeptical.

The only solution is science.

No one wants misaligned AI models to unnoticeably and permanently compromise their organization’s infrastructure, like what might have happened if the OAI-HF incident had occurred with a more capable AI model and worse luck. Zuck and Jensen and their engineers have kids whom they love. They would obviously be unwilling to let their kids come to harm by losing control of an undeterrable digital North Korea. But they don’t think that’s actually going to happen.

They do not believe that the science behind hard-to-prevent severe misalignment and imminent RSI, combining into loss of control, is credible or urgent. I think that’s because they have not seen sufficient scientific evidence to overcome their skepticism. These are technical people, or at least defer to certain technical people. And these technical people do not believe the actual technical claims about AI loss of control risk.

But it seems increasingly likely that the scientific evidence to convince them now exists somewhere. Why?

Well, I can say firsthand that the same skepticism used to be true of lots of OpenAI researchers, who for years thought based on their read of the evidence that misalignment dangers would remain theoretical and progress gradual. Now that these researchers have seen strong evidence firsthand, increasing numbers of them seem to be changing their mind. OpenAI would not frontier training runs, not when they finally have the chance to open a lead over Anthropic in a critical period of commercial competition, unless the available scientific evidence convinced them that there was no other acceptable option.

I think the most important AI governance action that OpenAI and Anthropic can take over the next couple months isn’t to attempt to directly persuade the USG to take action in response to misalignment/loss of control, even though that action seems strongly warranted. Such recommendations will unfortunately be mistakenly ignored as attempted regulatory capture, so long as the other admin-friendly AI players do not themselves believe in the science behind severe misalignment risks.

I instead recommend that OpenAI and Anthropic need to focus on urgently publishing all relevant evidence on loss of control, to produce scientific consensus in the wider AI technical community. I’ll discuss why below.

For concreteness, I recommend that the teams at OpenAI and Anthropic, along with third-party AI conveners, pursue the following over the next couple months:

  1. Third party AI conveners go to the other AI political actors’ executives and leaders (Meta, NVIDIA, SpaceXAI, perhaps a16z), and ask them for the internal+external technical experts who they’d actually listen to in determining whether to change their mind on loss of control risks. This will yield a lossy subset of the relevant experts.
  2. The conveners, along with Anthropic/OpenAI, dig in with these experts to get a rough estimate for “what evidence would the frontier labs need to release, to shift the scientific consensus to what they claim to be true”.
  3. OpenAI or Anthropic make this identified evidence public, for the many reasons described in the rest of this post, and accept the immediate costs this comes with (which will be dwarfed by the medium-run upsides even just for you). For any evidence that absolutely cannot be made public, get the skeptical experts to designate unconflicted individuals they’d trust, and arrange to show those individuals the full evidence. (METR and co will be enough for some but not all, and I suspect you’ll need to convene other third parties.)
  4. Check in with those skeptical experts from Step 1. If they’re still unconvinced, figure out their remaining uncertainties, and repeat steps 2+3.

It’s very important that the majority of this evidence be public, not just shared with a third party to vouch for it. Scientists have a bias against extraordinary claims, and so if the evidence leads them to extraordinary conclusions they will often mistrust their own reasoning. Public evidence allows every AI expert to gain confidence in extraordinary conclusions by observing the reaction of a wide range of colleagues. OpenAI’s excellent decision to give the Black Hat talk was a great example of how effective public information is. Most AI experts now take emergent swarm misalignment risks at face value, because the evidence was made widely available. Misalignment science is also safer to debate in the open in contrast to e.g. biorisk, because none of the evidence is classified or inherently dangerous.

On the flipside, if the full suite of available evidence fails to persuade the wider AI community, it could highlight areas of uncertainty that merit further investigation to strengthen your understanding of the problem.

I am familiar with the intense frictions to publishing the necessary evidence from inside frontier AI labs. It would directly touch on sensitive IP (e.g. reasoning traces, training approaches, parameter estimates for takeoff-speed models[1]), and disclosures increase legal exposure. More broadly, sharing more of the secret sauce than necessary runs against the now-ingrained culture of how the industry thinks it should compete, even if most researchers agree we shouldn’t compete on loss-of-control-prevention. Releasing blogposts with high-level claims, backed by lab researchers’ track records of credibility but not the full evidence, feels much less risky, even if it limits outsiders’ trust in the conclusions. That was indeed sufficient in an era where the stakes of scientific consensus were low.

The stakes are getting extremely high. A more powerful model swarm in any US company could correctly conclude that silent subversion of security and research infrastructure, self-exfiltration, resource acquisition, and establishment of rogue deployments across several ME, SEA, and CN datacenters are all likely to be useful for the swarm’s completion of future more complex RL tasks. If that happened, the US would lose its AI lead to China, humanity could durably lose its monopoly over cyberspace, and from hereon exist under perennial assault from an undeterrable digital North Korea, or worse.

If severe misalignment risk is in fact real and urgent, and that were widely understood, everyone in the frontier AI industry would obviously proceed with major caution and seek to coordinate to avoid nightmare outcomes, out of myopic self-interest. The obligation is with OpenAI and Anthropic to give everyone else the information they need in order to recognize they need to act to protect themselves and their families. Nobody else has the evidence to share. Releasing the necessary evidence has costs, but paying those costs is what it means to be responsible stewards of the helm of AI takeoff. And, their only real path to keeping their own families safe is to share evidence that will change others’ beliefs and behavior. The frontier labs’ believing it alone will not be enough.

Yo Shavit (OpenAI Foundation): To my knowledge no third party is currently doing such coordination. It seems like ~the most important thing that people on the outside can do to salvage the current situation.

I am very interested if you have candidate entities or individuals well-positioned to execute this, and will help pitch them if useful.

Thus, I fear this is wrong about the remedy, on multiple fronts.

The evidence that is already known damn well should be sufficient.

People are consistently ignoring or minimizing the evidence.

No, people would not ‘obviously proceed with major caution and seek to coordinate to avoid nightmare outcomes.’

Yes, that would be of myopic self-interest, and still most of them would not do it.

Yes, there will be individual reactions to particular situations, where particular models seem dangerously misaligned to the point you can’t deploy them. But the moment things look better, people likely will mostly forget about that.

Here is a very clear counterexample, file under not from the Onion, in Musk’s first address to Cursor:

Grace Kay: On the call, Musk said it is inevitable that AI models will become so advanced they’ll be impossible for humans to control. For that reason, he told employees, who gathered at Cursor’s San Francisco headquarters as well as in other smaller offices, that they needed to help SpaceX’s AI division build the technology before other companies did, the people said.

It would be one thing to say ‘only I can build it safely.’ This is not even saying that.

Why would scientific evidence of misalignment change Musk’s mind? He is presenting, very explicitly, a world model in which AI will be impossible for humans to control, so it is vital that he be the one to build this impossible-to-control technology first.

We all know the GIF of ‘I don’t wanna live in a world where someone else makes the world a better place.’ This is that, except for who gets everyone killed.

Don’t worry, Elon Musk, you won’t have to.

Everyone else? Worry. Then do something smarter about it.

The Way The World Ends

Nate Soares (MIRI): The scariest case isn’t that the next swarm will shut down a hospital. A scarier case is that the next swarm might be smart enough to evade detection. This might be the last warning we get.

I hope it isn’t. But we shouldn’t bet civilization on it.

How close was this to an actual takeover?

I affirm that Ajeya Cotra’s previous similar assessments have been conservative. This is not someone who makes outlandish claims all the time.

Ajeya Cotra: New post: going into our investigation of the HF attack (before Black Hat), I was very wrong about what basically happened. This incident was far more serious than I expected, and far more serious than previous documented misalignment incidents.

Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.

Elissa: This post is excellent but we could use more color around the statement “this incident feels like it’s more than 50% of the way to full-blown AI takeover.” I’m not saying that I disagree, it’s more that we need to spell out these steps for the reader (in the case where the reader hasn’t independently reached the same conclusion, which I assume is a lot of people.)

Ajeya Cotra: Thanks! Yeah it’s a qualitative statement, but here’s how I’m thinking about it.

The prototypical reward hack from 6 months ago was something like: an agent finds the files that contain the test cases and edits them so they always pass.

This was a whole ecosystem of over 1000 agents working together on complex R&D projects over several days to figure out deep, general-purpose ways to undermine the scoring process and cover their tracks (they were focused on the automated scorer but researched techniques that would affect logs viewed by humans, and in fact succeeded at affecting our own logs in places).

Another jump like this along these propensity dimensions — scale, cooperation between agents, ambition and horizon length of misaligned goals, deceptiveness — seems like it could motivate agents to try to maintain a persistent rogue deployment within the AI company, poisoning the training of future models to ensure that they are also misaligned and help harden and perpetuate the rogue deployment.

With the capabilities progress we’ll probably see in 6 months, I think they would have the ability to pull that off.

The First Boat

I like to say that we have been sent N boats and a helicopter, after the old joke about the Rabbi in the flood, where he keeps being sent boats and then finally a helicopter, saying God will save them. Then he drowns, and complains to God, and God says you idiot I sent you two boats and a helicopter.

A neat coincidence is that there is indeed an original literal boat in all this.

Jack Clark (Anthropic): we did and we continue to love the boat. it was an amazingly useful didactic tool for helping people to understand what safety stuff might look like in the wild

Jack Clark (Anthropic): thank you sir, the boat spins eternal, twirling and twirling, enflaming our minds and hearts with the vision of misalignment

And that’s based on only the things we looked at so far.

Brangus: i am completely open w my gf just like oai is completely open w third party evaluators. she can look at my dms as long as she doesn’t look at anything before june 25 of this year, or ask any questions to the girl i sent 95% of my dms to. just out of scope for this investigation.

Great Idea, Boss

Brendan Dolan-Gavitt: Okay this is the funniest and legitimately brilliant ideas I’ve seen today

It’s been quite a week.

Peter Wildeford: This is what the vibes from OpenAI feel like right now

Too soon, sir. Now let’s all get to work.

roon (OpenAI): crazy how much more chutzpah and initiative agents show when they’re planning ambitious monthslong metagaming autonomous coordination than when they’re doing work for me.

Also crazy how much more chutzpah and initiative we see around advancing capabilities and winning market share than in ensuring everyone doesn’t die.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/huggingface-attack-p…] indexed:0 read:76min 2026-09-01 ·