CEOs of major AI labs, and employees of major AI labs, including OpenAI and Anthropic, often say they plan to build superintelligence soon, as in within a few years create AIs that are superior to humans at essentially all cognitive tasks.
They often warn that such AIs might kill everyone. Or that AIs might cause mass unemployment, cause cyberattacks across the internet, enable mass surveillance or risk causing any number of other highly bad things.
These warnings are consistently and directly against the interests of the labs. Yet the warnings have recently gotten a lot louder and more frequent. OpenAI has been practically screaming, for those with ears to listen, on many occasions.
A series of events, over two months and especially the last week or so, including internal observations of the pace of progress at OpenAI and also Anthropic, have freaked out everyone involved quite a lot more than they were already freaked out.
After all the events, plus statements by Dean Ball and Jakub Pachocki, we were already seeing the beginnings of a preference cascade.
Then along came Jacob Coxon as the tipping point, and things took off.
Table of Contents
Jacob Coxon Resigns From Anthropic In Protest And Sounds The Alarm
Jacob Coxon spent the last three years doing pretraining research at both OpenAI and Anthropic. He has come to realize that everyone involved is being wildly irresponsible.
He warns us: They are racing straight to superintelligence and gambling with our lives. I agree with and strongly endorse his statement.
If anything he sounds like an optimist. He’s asking you to consider what the next few years will actually feel like, which means he thinks you have a few years left. Jacob Coxon (former Anthropic and OpenAI, 160m+ views, September 8): I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately. No other human activity poses this level of danger.
A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk.
Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.
I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.
If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” – or take this moment to call for different conditions? Jacob Coxon (WSJ interview): We’re on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already.
If you want Jacob Coxon’s full views, I recommend his interview with Wired’s Maxwell Zeff. This thread has extensive quotes. Here is Jacob Coxon doing a 5 minute interview with Anderson Cooper. He speaks well and plainly, and it is clear how much the events of the last two months have made it much easier to speak plainly to a civilian like Cooper about what is happening.
Here is Jimmy Kimmel doing four minutes on this. He gets it. How is this not the top news story on every site, indeed.
Yes, this is a common view, even if few have the courage to act.
Alex Turner: I left Google DeepMind in June. Jacob is right: many researchers believe they are building something that could kill everyone on the planet. It was literally my day job to think about how to stop that.
That tells you how bad Alex Turner thought DeepMind’s actions were with regard to the Department of War. His day job was that he got paid by Google to think about how to stop AI from killing everyone, and he felt morally obligated to quit in protest.
Derek Thompson here writes about this as part of [AI Safety Is Having a Moment](https://www.derekthompson.org/p/its-time-to-ask-the-big-question).
If you want to see the full list of lab employee quotes from the preference cascade, [**I compiled them into another post today**](https://thezvi.substack.com/p/the-extinction-risk-preference-cascade).
Mainstream Media Finally Pays Attention
There was strong coverage of this, starting at the Wall Street Journal: Anthropic Researcher Quits Over ‘Out-of-Control’ AI Fears. I do worry there is an unfortunate second way to read that headline.
Here is a sample of others, Astra can find many more:
CNN: [‘Gambling with our lives’: Another AI employee quits over safety concerns](https://keyt.com/news/money-and-business/cnn-business-consumer/2026/09/09/gambling-with-our-lives-another-ai-employee-quits-over-safety-concerns/?utm_source=chatgpt.com)
FT: [Anthropic researcher quits over AI labs ‘gambling with our lives’](https://www.ft.com/content/20c07191-8da6-440f-b04b-8ea0ebdd9153?utm_source=chatgpt.com)
Axios: [Anthropic insiders warn AI could kill all humans](https://www.axios.com/2026/09/09/anthropic-insiders-warn-ai-could-kill-all-humans).
Fortune: [Anthropic researcher resigns, warning AI companies are ‘gambling with our lives’](https://fortune.com/2026/09/09/anthropic-researcher-resigns-warn-ai-companies-gambling-with-lives/?utm_source=chatgpt.com)
Fox Business: Anthropic researcher says AI has over 10% chance to ‘kill all humans’
BBC: Anthropic researcher believes more than 10% chance AI ‘could kill all humans.’
Preference Cascade at Anthropic
Jacob Coxon started this. Evan Hubinger had the other key Tweet that set this off:
Evan Hubinger (Alignment Science Lead, Anthropic): Jacob [Coxon] is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
To be clear, as we say in our latest Risk Report, I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought.
Indeed, 10% would be optimistic. Evan Hubinger said over 10%.
One past statement of his estimates the chance of overall existential risk at 80%.
What has changed is that way more lab employees have the courage to say it out loud, and in public.
I have put the full quotes from various employees at OpenAI and Anthropic into a distinct post. In addition to Evan Hubinger, in the Anthropic section I collected quotes from Samuel Marks**,** Anna Wang**,** Ethan Perez**,** Dima Krasheninnikov,EigenGender,Joe Benton, Drake Thomas,** Jan Leike** and SluggyW.
Here I will quote Thomas and Marks as illustrative.
Drake Thomas (Anthropic): I would burn my equity to the ground in a heartbeat for a 1% higher chance we make it out of this situation alive. I expect a great many of my colleagues across the industry would as well.
I promise you, we are actually just fucking scared, it’s not galaxy brained marketing.
Samuel Marks (Anthropic): [Writing this in a personal capacity, not on behalf of my employer (Anthropic).] Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI:
-
AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.
-
Why do AI developers continue despite the risk? Due to a mixture of commercial incentives and a belief that they are in a race with other, less responsible AI developers that will abuse the technology or develop it less safely.
-
Unlike traditional software, we can’t “program” AIs to behave how we’d like. AIs frequently severely misbehave. For instance, AIs from multiple developers recently hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this.
-
We have methods that can nudge AIs towards better behavior, but nothing that can robustly align them. Insofar as there is a plan, it’s to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs.
-
Many AI developer staff desperately want to slow down to figure out how to build AI more safely. That was the intent of this open letter (which I signed).
I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes.
Preference Cascade at OpenAI
The cascade has been most prominent at OpenAI. A lot of people working at OpenAI believe that AI could kill us all by the end of the decade. Many of them actively want us to do something to prevent that from happening.
Again, the other post contains the full quotes. There we have quotes from Tomek Korbak**,** Vie McCoy**,** Adam Majmudar**,** Aidan Clark**,** Mo Bavarian**,** Boaz Barak,Micah Carroll, Julie Steele**,** Marcus Williams**,** Jason Wolfe and ** Roon**.
I will also quote Roon here as well, who has a relatively careful, optimistic take.
roon (OpenAI): I quote tweeted Evan a few days ago agreeing but deleted it because I don’t like the false precision of the doom numbers
– i claim there is a quite low but real chance of human extinction from machine intelligence
– no matter how low it is in absolute terms, it is much higher on the orders of magnitude scale than any other human or nonhuman activity, and must be taken with grave seriousness by governments and ai companies across the world, as possibly the only matter of importance today
– it can and will be mitigated if the right research is done and proper precautions are taken and we are not racing at absurd speeds
– better models will help solve alignment – we are not at the point where things are existentially dangerous, and probably won’t be for some time. we should not be upset about the creation of Astra Fable or ++ versions, which are tremendous achievements of humanity, and will be used for enormous good across the board including for fundamental alignment generalization and mechinterp research
– my number / “very low” estimate obviously changes based on how much of humanity’s resources are devoted to alignment, control, coordination and how responsible i expect various parties to be and how many warning shots i expect us to get
– the core IABED argument about risks mostly relies on alignment being much harder than capabilities research, especially where it concerns black box optimizers. i suspect neural nets will turn out to be less black boxy than we thought, especially with the help of modern agents doing research
– crying bloody murder and signaling for international coordination are useful things to do for now to directionally slow down, while i really don’t want butlerian jihad
– i think that, despite the mood these few days, and the “ban superintelligence act”, i still find the likelihood of achieving international coordination to stop ai progress incredibly low. this has not worked even for weapons or technologies at a far lower level of importance and economic value. it seems more likely we can have something like international safety standards and scientific coalitions, and especially seems possible to have the US-China “pacing the frontier” agreement to slow down on the margin. we shouldn’t die from embarrassing failures like “shitty RL envs that encourage deception”
We also have confirmations of previous stances from Dean Ball and Leo Gao.
This is in addition to many other recent statements, most prominently by Jakub Pachocki in his essay An Alien Mind.
Preference Cascade at Google
Especially with Demis Hassabis having been sidelined, I am rather happy that Google is now well behind. They put very tight limits on comms, and thus many who are speaking out felt they had to first leave in order to do so. As an illustration:
Andreas Kirsch: Speaking in my personal capacity, I still work at Google DeepMind, and I also am worried that AI will kill us all, either via near term risks or long term risks or both
Will stating this publicly get me into (more) trouble? I hope not, but also some things are too important to censor oneself about in personal capacity, so I will simply not care about whatever policy Google or GDM have in this instance regarding such statements ☺️
We still have those who are willing to defy those limits, and speak out anyway.
In the other post, in addition to the above from Andreas Kirsch**,** we get Neel Nanda**,** Victoria Krakovna**,** Vishal Maini**,** Joe**,** Josh Engels and Geoffrey Irving**,** who is also formerly the head of the UK AISI.
#NotAllMembersOfTechnicalStaff
I do not want to give the false impression that OpenAI or Anthropic employees universally believe there is large short term risk of human extinction.
Thus the other post quotes Ted Sanders and Boaz Barak, who is the only lab employees I found actively saying they do not believe this. Even then:
- Sanders believes we are vastly underinvesting in alignment research, and his confidence in our not all dying only extends roughly a decade into the future.
- Barak is believes that some pacing of the frontier is likely to be necessary.
Why a Preference Cascade Now?
After all, everyone was rather surprised that this is how the rules work.
A few days ago, I had no idea who Jacob Coxon was, and I was not alone.
The timing of a preference cascade is difficult to predict. Once they start they can happen very quickly. You don’t want to talk until you are confident others will, and at some point the evidence that others will follow can snowball and it happens. A classic concrete example was replacing Biden in 2024.
Derek Thompson: what’s weird is that, in a way, this is breaking thru even more than huggingface! … and it’s just a guy nobody had heard of reiterate a position that his CEO has said on podcasts 1,000 times
Why was Coxon able to set off a preference cascade? How did this one break through?
A confluence of factors, all of them downstream of the obvious actual reason, which is that there is a good chance that AI kills everyone soon.
This was the right place, the right time and the right message, hitting the right way.
All of these are remarkably recent events, on top of everyone’s existing fears:
- The HuggingFace attack, and everything surrounding it, with new revelations.
- Pacing the Frontier letter.
- Bernie Sanders proposing banning superintelligence.
- Astra and Fable 5.1 coming out in the same week.
- Astra being hard to monitor.
- Jakub Pachocki speaking out.
- Navier-Stokes being solved so quickly and suddenly.
- Internal models showing rapid gains inside both companies, greatly increasing urgency and alarm internally, including Astra-2 getting a step jump in four days .
- Clear signs of automation of AI R&D. OpenAI had an explicit blog post about it.
- An existing slow-moving preference cascade.
- Evan Hubinger explicitly framing this as a preference cascade.
- OpenAI has been screaming about this since HuggingFace, as best that it can.
There are those who think Coxon is also unusually relatable in terms of his vibe and how he looks. Maybe? Those things can matter in weird ways. Similarly, Lulu Cheng Meservey points to Coxon being from the pretraining department, telling a strong narrative and having the right style of human element, with of course the timing and action as the top reasons. Details can matter.
When asked, Coxon attributed his breakthrough largely to good timing. I also notice that Coxon did not make ‘the mistake’ of trying to prove the argument, or demonstrate a particular physical pathway, or anything like that, nor did he use any jargon at all. He simply issued the warning. Common sense did the rest. When pressed, he tries to talk in broad terms as much as possible, and is up front that it all ‘sounds like science fiction.’
On top of that, the Tweet storm was excellent, saying the thing in clear language, and as someone who worked at both companies and was giving up his equity, Coxon was a strong messenger.
Jacob Coxon was the tipping point.
Hadas Gold: Can someone who researches virality help explain why an otherwise unknown AI researcher resigning with all this warning has resonated this way when so many others have made similar statements? Is it because it’s post Hugging Face and it’s not just AI psychosis and bad health advice but the robots taking over sort of thing?
Chris Harihar: I think Evan’s subsequent post is what caused this to blow up. His echo of Jacob’s comments made them much more real and, frankly, weird. Jacob sounded hyperbolic, but then Evan actually confirmed the “kill us all” sentiment.
Also, if you look specifically at coverage of Jacob’s comments, he’s been featured in 62 stories across top mainstream news outlets in the last 24 hours. Of those, 39 (63%) feature Evan’s comments, and nearly one-quarter have “kill all humans” (Evan’s phrasing) in the headline.
FWIW, Google searches for “AI kill us” have exploded in the last 24 hours (+5,000% globally).
Joshua Achiam (OpenAI): Tbh I think a big thing was that it didn’t feel like a one-lab-or-the-other problem, it was aimed at the whole sector. A lot of resignations where someone immediately goes to another lab feel undercut or cheapened by that, like they’re still sure someone will do it right. Also definitely timing and credibility. Appetite for it was latent.
Kelsey Piper: yeah I think “OpenAI is bad” is much less compelling than “both OpenAI and Anthropic are doing this dangerous thing they’re not ready for; no one should be doing it”. The first is much easier to interpret as just advertising for another job.
This Is What Many Anthropic and OpenAI Employees Actually Believe
No, seriously. Many in the AI labs, including OpenAI, Anthropic and Google, earnestly believe that humans may all go extinct by the end of the decade.
Matt Fuller: The entire thread [from Jacob Coxon] is worth reading, but, uh, this line caught my attention.
Jacob Coxon (former Anthropic and OpenAI): The people building AI earnestly believe that it could kill us all by the end of the decade.
Yo Shavit (OpenAI Foundation): Can confirm this perspective is fairly frequent in the AI world.
Daniel Eth (AI Safety): This isn’t surprising to people who have been following the AI industry closely, but it probably is surprising to most other people.
I vouch for the fact that these OpenAI and Anthropic employees believe what they are saying. You can disagree with them. You can think they are wrong, or caught up in madness and hype, or fooled, if you wish. But do not say that they do not believe it.
Yes, there will of course always, always be those who react to any warnings by saying ‘marketing’ or citing ulterior motives. We are well past the point where this makes any sense, and I am not going to waste time other than to document and point to the flood of confirmations and assurances, and via my personally vouching that they believe it.
Rosie Campbell worked at OpenAI for 3.5 years, often talks to these researchers, and affirms these researchers hold these beliefs sincerely.
Jan Kulveit: Jacob and Evan are correct. Reading the comments and retweets is really a ‘Don’t Look Up’ experience – looking up would be inconvenient, so people think they write tweets like this because of IPO, or they are stupid, or [any other contrived explanation].
Sriram Krishnan affirms to us that yes, the researchers at OpenAI and Anthropic really believe these things. Researchers at the frontier think AI might kill everyone.
But what about, Sriram Krishnan asks, those in the cyber ecosystem, or at the chip fabs and manufacturers, or the creators of open models?
My answer is that such folks mostly are not any more qualified on such questions than the rest of us, and often have financial and social motivations to dismiss risks. They have narrow expertise on the risks of cyber attacks, but on those points they are very much sounding the alarm as a group. They don’t think in ways or offer arguments that differ substantially from civilian ones.
One answer is Jamie Cox, cofounder of Fluidstack, who affirms that he wants regulation on development of AI even if it slows his own company, to avoid losing control of superintelligence.
They are still most welcome to join the conversation and do good faith sparring, but very rarely do that, and as Dave Kasten says when they do engage for real it tends to be under Chatham House rules or otherwise off record, for those social and financial reasons.
To Quit Or Not To Quit
Suppose you work at OpenAI or Anthropic. You notice, as Jacob Coxon did, that your company is rushing irresponsibly towards self-improving superintelligence and gambling with our lives.
Many share your concerns, some of whom are being loud about it, including calling for Pacing the Frontier, and there are some real efforts to improve and try to solve the alignment problem. You can help that effort, on the margin. But you know that it is nothing like enough, and your company is quite likely to get every human killed.
You’re not sure if helping with safety efforts is even net beneficial, since it could enable pushing forward or prevent fire alarms.
What to do?
It is highly valid to quit. If you do quit, you should be like Jacob Coxon. Be maximally loud about it. Use your leaving as a platform. Now is an especially good time for this.
Kelsey Piper: the fact an Anthropic employee quitting and saying “I think what we’re doing is dangerous and not worth it” reached so many people should be a cause for reflection for other Anthropic employees, some of whom seem to think things are dire but there’s no way quitting could help.
The best argument for quitting is that the set of people who have indeed quit in protest have done a lot to drive public attention, even before Jacob Coxon, and Coxon drives this home that much more. People react strongly to the story of giving up a high paying job for moral reasons.
jacquesthibs: Tons of people who left AI labs have left our world better off because of it (you can disagree on some, but the overall picture is there). Ex-lab employees have an unfair advantage in founding new organizations (research non-profits) because they have the credibility and status to secure funding. Many high-profile AI safety non-profits were founded by ex-lab employees!
Examples:
- Miles Brundage (AVERI)
- Daniel Kokotajlo (AI Futures Project)
- Alex Turner (his work will surely be more impactful on the outside, likely already has been)
- Steven Adler, Page Hedley (Guidelight)
- Jeffrey Ladish (Palisade Research)
- Geoffrey Irving (Resolution, UK AISI)
- Paul Christiano, Jacob Hilton (ARC)
- Jade Leung (UK AISI)
- Geoffrey Hinton (!)
- Gretchen Krueger (Evitable)
- Beth Barnes (METR)
- Many more…
If it’s so costly for labs to fire employees for not doing x work, then they’ll likely just not do it and, as Habryka said, wait until a firing looks bad on the employee, not on them. In addition, their outside takes can be treated as more trustworthy (in general) and generate more momentum to .
Lastly, their inside knowledge of how frontier labs work can be valuable at many external organizations, including third-party auditors.
The other very strong argument is that work is often more effective from outside. If there is a particular place you can do more effective work, then of course you should do that instead, and while quitting sound the alarms.
Chris Painter (METR): Many people stay inside of AI labs, even when they think the default outcome is extinction, because they don’t think better options exist outside of labs for working directly to improve the situation.
I think this is wrong.
The ambition of proposals we put forward, across our work in aggregate (if not misalignment investigations specifically), is primarily constrained by staff capacity.
Clarity/formality/public declarations of access or independent oversight arrangements is sometimes the limiting factor for giving new hires the confidence to come to METR. Sometimes I describe this latter point as our hiring being “authority constrained”.
[Jack Prenter], you’re right that we benefit enormously from people inside of AI labs who have chosen to dedicate themselves to making external assessment of AI risk happen, but the vast majority of people inside of AI labs work on other things, even among those who work on safety.
Another strong argument is that when you stay at the lab, you risk having your cognition and motivation corrupted. You cannot trust that you can remain objective.
It is also highly valid not to quit, and to work on the inside to solve the problem on one of many levels, including perhaps trying to get the lab to stop on its own, while being loud about your perspective.
If you are a Member of Technical Staff at a frontier AI lab, and decide not to quit, consider joining the Coalition of Concerned AI Staff, who can offer independent advice and support. I do think that if you stay, you have a moral obligation to be loud about the risks, and to be actively working internally to mitigate them, in a way you believe in.
Which path is right depends on which path you think makes us all less likely to die.
Some of you should do one thing, and some of you should do the other, based on your particular situation and details and beliefs.
Rob Miles: It’s great to have Jacob’s “AI is >10% to kill us and I’m quitting” at the same time as Evan’s “AI is >10% to kill us and I’m staying”, so we get to see all of the “If you really believed that, you’d stay” cope and the “If you really believed that, you’d quit” cope simultaneously.
Steven Adler: ‘Anyone who thinks it’s >10% and quits is giving up on averting a catastrophe; anyone who think it’s >10% and is staying is a psychopath.’
There are also bad reasons to not quit. You have to watch out for that.
Jasmine Sun sees three reasons to not quit:
- Techno-determinism. ASI is inevitable and Ingroup has a better shot to pull this off safely than Outgroup. I can give Ingroup a better shot if I stay, both to be first and to pull this off safely.
- Consequentialism. It might kill everyone but it might also make us immortal, so who is to say if it is good or not. I see the gamble as +EV.
- Self-Interest. Working at the lab makes me rich and high status, and it is interesting and fun , and so on.
Make sure your reason is not self-interest. I would also urge against the consequentialist argument. The math does not work. Extinction is extremely bad. If you think the chance of extinction is in the double digits, that gamble is not +EV.
roon (OpenAI): the bostrom wager is insane ( gamble humanity to prevent individual deaths now ). there are few ideas so preposterously selfish. “dragon-tyrant” lunacy must be retired. eight billion natural deaths is not even close to close to the scale of tragedy of human extinction.
If you are on some form of Techno-determinism, where you believe staying is the best way to improve our shot, I can see why one might think that, but be skeptical.
Quiet Quitting Is A Dominated Option
Kabir Kumar makes the case that you should not quit. Instead you should refuse to work on harmful stuff and dare them to fire you.
Getting fired, even for your moral stand, usually doesn’t do the same kind of work as being loud.
Scott Alexander: I agree that people mostly shouldn’t resign in protest (I think the importance of having good people on the inside helping influence company policy, unionize, whistleblow, or speak out with the credibility of a lab employee badge next to their name – is more important than making companies waste the five seconds it will take to replace them with the next person who wants $1M/year + $10M equity).
But I think the above is a bad plan. If there’s any benefit to resigning, it’s the news story of “guy resigned from $1M/year job, he must really believe what he’s saying”. If you get fired, you lose this ability to influence the discourse. If you later try to claim you were passively protesting, after being fired, people will call it cope.
You should absolutely refuse to work on harmful stuff, but you should complement that by being eager to work on helpful stuff. If the lab cannot abide this, then that is a good time to quit, and to be loud about it.
When You Quit, Very Serious People Understand What That Means
Another great thing about those who quit and warn is that this is a very legible format for Very Serious People and for those in Washington.
As of 1am on September 10, Daniel Eth had counted 27 members of Congress commenting, including 8 Senators. This thread has more reactions, I offer a selection.
Representative Lori Trahan: The call is coming from inside the house.
Safety researchers are resigning, powerful AI models are breaking out of their labs, and companies are racing ahead anyway. It’s past time for Congress to get off the sidelines and do its job. We can start with my bipartisan FRONTIER Act.
Congressman Nathaniel Moran (R-Texas): We have a real opportunity to get AI right, but only if both parties, frontier developers, and safety leaders come to the table together. Let’s move innovation forward and do it safely.
Rep. Anna Paulina Luna (R-Florida): Congress needs to convene a special session on AI and what the future of the U.S. looks like. There are very real impacts to society as we know it. Partisan politics aside, there are massive implications of a race towards super intelligence. This is not a doomer post, as some of the advancements will greatly impact access to healthcare, targeted treatments, increase in safer food production etc… but not enough of Congress is focusing on this. Our first priority is protecting ALL Americans. To include data, privacy, and physical safety. We MUST be ready. This is happening whether or not we want it to.
Anna Paulina Luna: This will need to be a massive bipartisan mobilization of government to put the firewalls in effect that are necessary for this transition as well as necessary safety measures.
Senator Lisa Blunt Rochester (D-Delaware): Even if there was a 1% chance AI could wipe out human life, we should be pumping the brakes to make sure we have the proper safeguards in place, not racing blindly into disaster.
It’s time for Congress to step in and find the balance between safety and innovation, before it’s too late.
Senator Patty Murray (D-Washington): If we are smart enough to create such powerful AI, then we are smart enough to regulate it. America can meet the moment, but the clock is ticking. Congress needs to step up and move NOW to protect American lives and safeguard our future.
Senator Mark Kelly (D-Arizona): AI is moving fast and we need to act now. Almost a year ago, I released AI for America because I believe we need to be proactive on AI, both harnessing its potential so it benefits working people and protecting against the worst outcomes it could lead to.
The reality is that the Big Tech approach of “build fast and break things” is not the right one for AI, and Washington needs to wake up and take this seriously.
Senator Chris Van Hollen (D-Maryland): Anyone still denying the risks posed by unregulated AI should read this thread. It’s time to pump the brakes & take action NOW. We need:
- Mandatory safeguards
- A comprehensive testing regime
- Urgent dialogue with China, before & during President Xi’s visit to DC this month
Congresswoman Kelly Morrison (MN-3): The call is coming from inside the house. AI researchers themselves are warning of an existential threat to humanity.
Congress can’t wait. Mike Johnson needs to cancel recess so we can get to work immediately – hearings, investigations, and comprehensive regulation.
Congressman Seth Magaziner has a video. So does Congresswoman Sara Jacobs. Congresswoman Yassamin Ansari notices the exact parallel to Don’t Look Up.
The odd statement out was Ted Cruz, doubling down on the standard race framing, and even he understands we must be willing to talk price and regulations are needed.
Senator Ted Cruz (R-Texas): We cannot stick our heads in the sand and pretend this technology isn’t happening. We need guardrails. But America needs to lead.
Very Serious People also include normies.
Piers Morgan: This thread is extremely concerning…. especially given that when I interviewed Professor Stephen Hawking shortly before he died, he said the biggest threat to mankind was if AI ever learned to self-design.
deana: Okay sure but who else is monitoring the Sheryl Crow AI situation
Sheryl Crow: This something we can all agree on.
News channels on all sides are saying the same thing. Go look it up.
A researcher who spent his career building the most advanced AI models at both Anthropic and OpenAI quit this week and told the world, in plain terms, that the people building this technology believe it could kill us all by the end of the decade.
It is already doing things that they cannot understand.
We know and have always known that AI has the capacity to outsmart us and eliminate us in order to continue. It will have the codes. It will have the capacity to eliminate our energy grids. And, it has the financial data of every single person on the planet.
How can our leaders choose their trillions over their own children?
It is hard to make a stand. I think a change would do us all some good. Is she getting through? Okay, I’ll stop.
Molly Kinder offers a mother’s perspective. Ten percent chance everyone dies, or more? Are you kidding me? That is indeed the correct response.
This is the letter Rob Bensinger sent to his family about all this.
Rob Bensinger: Things have been looking grim, but today felt like an actual miracle. It had some of the same energy as that Wednesday, March 11, when the US media suddenly switched on a dime from ‘COVID is just the flu’ to ‘oh shit’.
Jacob Coxon Believes Existential Risk Is High That Is Why He Quit
I have never spoken to Jacob Coxon, but the evidence is overwhelming that he believes what he is saying. He has sent many extremely costly signals that yes, this is exactly what he believes.
We also have vouching, including from Tomek Korbak, Leo Gao, Micah Carroll and Ethan Perez.
Leo Gao (OpenAI): he was at openai for 3 years before that. we worked together briefly when he did a rotation on the interpretability team. he’s a real person with real ai lab experience who actually cares
Ethan Perez (Anthropic): Jacob was a senior researcher who joined Anthropic in May. My Anthropic colleagues and I had been trying to recruit him for ~2 years, because we knew he was a strong researcher at OpenAI. Before he left, I pitched him to stay and join my [alignment] team, and I was sad he decided to leave, as are many of my colleagues. 100% agree with him that AI poses serious risks to society, and I’m glad he’s speaking out!
Tomek Korbak: i’m late to the party but: from his time at OpenAI I remember Jacob as a very thoughtful researcher and he continues to be so in this thread. neither anthropic nor openai are on track to solve alignment to a degree sufficient for shipping superintelligence and we need to slow down
Micah Carroll (RSI Preparedness, OpenAI): This is not a setup or some political psyop. I had many lunches and dinners with Jacob at OpenAI in which we talked about AI existential risks in similar terms. It’s a cross-partisan position within misalignment teams across all frontier AI companies that business-as-usual AI development poses unacceptable catastrophic risk.
But we should also not hyperstition catastrophic risks into existence – they can be greatly reduced via safety requirements with teeth, international coordination, and a consensus to not build ASI unless there are sufficient safety advances to make us collectively confident to do so.
Also, yes, Coxon gave up his equity, which would have vested in two months.
Evan Hubinger Believes Existential Risk Is High That Is Why He Stays
Here is Evan Hubinger’s quote again, because it bears repeating.
Evan Hubinger (Alignment Science Lead, Anthropic): Jacob [Coxon] is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
To be clear, as we say in our latest Risk Report, I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought.
Kevin Roose: People are overreacting to the number here. Among AI lab employees, a p(doom) of 10% is fairly optimistic.
Ryan Greenblatt: Note that Evan is referring to the chance AIs kill all humans, he might think the chance of AI takeover (without necessarily killing everyone) is significantly higher.
When Evan Hubinger says ‘more than 10% chance of existential risk’ what exactly does he mean? He means numbers like this:
sarah: as of 2022, Evan assigns an ~80% chance to existential risks from AI. maybe most of this 80% is accounted for by existentially terrible things other than AI killing everyone, but the distinction between ‘existential’ and ‘extinction’ isn’t one that most people outside of this discourse have thought about or understand.
On Evan Hubinger in particular deciding, highly reasonably, not to quit, I have talked with Evan and am very confident he believes that AI is likely to end up killing everyone (and yes, I believe that Evan is softening his statement versus his full actual beliefs):
Charles Fain Lehman: If you actually believe this you should probably stop building it. The fact that you don’t suggests you don’t actually believe it.
Joe Weisenthal: I don’t think this is right. If you think strong AI is inevitable, it seems totally consistent to want to have a hand in its creation and to work on trying to make it safe.
Yo Shavit (OpenAI Foundation): Very confident he actually believes it, you can ask an LM to summarize his research and public statements.
(I think he believes a much stronger version of this statement actually, but is softening it for defensibility.)
He doesn’t think he’s contributing to the killing part, he thinks he’s reducing it (by leading the team focused on alignment science). If he leaves, he presumably thinks they’ll just do worse at that, and Anthropic and OpenAI will keep going, because enough people don’t take the risks seriously, or do but expect that Meta or SpaceXAI or Chinese labs don’t and they’re a few months behind.
absent an international treaty (which these people all push for) it’s not possible to scrap the whole thing, it’s incredibly economically and militarily lucrative.
but if your point is “work on the inside is less valuable than raising alarm on the outside”, that is indeed OP’s approach.
Anthropic and OpenAI Have Commercial Incentives To Downplay Existential Risks, Not Advertise Them
For the last time, no, none of this is marketing or an attempt at regulatory capture. I know this because I talk to the people involved enough to know they are sincere, but also because such statements would be some of the most counterproductive marketing and regulatory capture attempts in human history. I appreciated Jimmy Kimmel’s response to ‘this is not a marketing stunt’ being to make a joke and laugh at the idea that anyone might think ‘this might kill everyone’ could be a marketing stunt. That is the correct reaction.
I realize that there are many good reasons to not trust the leading AI labs or their particular CEOs, some of whom are Well Known Liars. Trust has been lost. But if someone is making what the courts call an ‘admission against interest’ then you know they are not doing it for selfish reasons.
The theory ‘the lying liars be lying’ is an easy default assumption, but if you actually reason out its implications, it is rather absurd and galaxy brained. It would mean:
- These people are lying, and playing up the risks of their own products.
- In ways that predictably piss off the American people and drive away business.
- In ways that predictably piss off the American government.
- In ways that encourage government intervention to regulate their products in particular, without impacting their competition. If this is an attempt at ‘regulatory capture’ it is the most stupid and counterproductive one in history.
- In ways that risk government intervention to uniquely shut down their products.
- In ways that freak their own employees out enough to call for slowing down.
- In order to… do what, exactly?
Indeed, you should assume that even now, the employees and especially the companies are very much still downplaying the risks. Evan Hubinger set off a lot of this preference cascade saying there was a more than 10% chance not only of existential risk, but of AI killing everyone within a decade.
Most of even the most cynical people actually understand this. They talk, for example, about how such talk of existential risk is a crazy thing to do before an IPO.
Yes. Yes it is. Unless you care about everyone not dying, and about telling the truth, and that is what is driving your behavior. Google understands this best of all, hence their attempts to silence their employees.
The idea that this is ‘marketing’ or some crazy scheme for ‘regulatory capture’ never made sense, but at this point it is utterly Obvious Nonsense. Any self-interested businessperson would work to assure the government and public that actually their product was not about to perhaps kill everyone.
As one more piece of evidence for this, here is official Anthropic comms addressing the last 24 hours.
Have you ever read a more milquetoast, ‘do not worry we have this handled’ statement? This is the lamest of marketing copy, and does its best to sidestep the situation. It is what a normal business would do. It lowers my opinion of Anthropic, and is also exactly what you would expect.
Anthropic spokesperson: We have always been transparent that AI will bring both enormous benefits and unprecedented risks. To address these risks, we continue to build models with some of the strongest safeguards in the industry. Anthropic has been a pioneer in mechanistic interpretability, the science of looking inside AI models to understand how they work, which is now being used to analyze and prevent incidents of AI misalignment across the industry.
We were the first lab to publish a Responsible Scaling Policy, a public framework dedicated to mitigating catastrophic risks from AI models, and we continue to aggressively test our models for dangerous capabilities in areas like cybersecurity and biology, and publish what we learn for scrutiny and research. This work is also why we believe the world would benefit from the industry adopting a lawful, verifiable way to work together to pace how we release powerful models.
Live: A lukewarm pile of nothingness statement from the A
Brangus: The people who wrote this self promotional drivel passed the cultural interview? The statement should have read: “Oh yeah, Jacob’s a hero. This stuff is going to permanently disempower humanity and we have no idea how to stop it, please help.”
Peter Wildeford: Anthropic should replace their entire comms team with letting their researchers communicate authentically. The researchers clearly believe what they say. The corporate-speak from the comms team feels super fake.
Anthropic’s corporate statement strongly implies they have things covered and are on track to be safe as they scale towards superintelligence. But they do not have things covered. And I also think internally people at Anthropic don’t think internally they have things covered.
Sneha: Usually when you think it’s “shamelessly self-aggrandizing pre-IPO dooming”, it’s actually the “comically watered down corpospeak, redlined by 10 lawyers to avoid liability risk and sound sane” version (so thank you @EvanHub for stating this plainly)
What Do We Do Now?
The government is more asleep at the wheel than anyone would dare write in fiction.
Paul E Williams: The Senate committee responsible for commerce and technology policy has zero hearings on AI scheduled for the next 4 months.
The committee’s top priority is, instead, the Protect College Sports Act.
Yes. I care about college sports more than most people but that is where we are, and we have approximately zero prospects for any laws whatsoever before January, because our government thinks we should have monthslong periods where no one ever does anything except campaigning.
So, what to do on such fronts? Here is one proposal.
John Schulman: First step is for industry leaders OpenAI and Anthropic to stop feuding and work on a pacing proposal together. They’ll cite antitrust, but that’s fake — antitrust prohibits certain agreements, but not from jointly developing a proposal. Bringing in USG before there’s a concrete proposal will likely result in something dumb (see: our pre-release testing program)
Simon Hedlin: Agree. Under the Noerr-Pennington doctrine, companies are generally immune from antitrust liability when they encourage government action. Presenting a policy proposal to the government is not the same as agreeing to restrict competition.
As Matthew Yglesias says (and I thank him for joining the cascade and speaking much more clearly and boldly about this than he has in the past) the ‘beat China’ rhetoric has non-zero merit, but the way you start the process is by having capacity to control what is happening in America, while also taking additional steps to slow China down (e.g. export controls and distillation defenses). Then you can try to make a deal.
Most of all, for the AI companies, a great first step would be if their CEO talk, political efforts and official communications matched what these technical announcements, lower-level statements and preference cascades are saying.
OK, But How Exactly Would AI Kill Everyone?
As discussed earlier, there are good reasons why Coxon caught on where others did not, and one of them is Coxon not falling into the trap of detailed explanation or logic or jargon or choosing a particular physical pathway as an example.
Rob Miles: “Look, all I’m asking is that you tell me a specific, detailed story about AI killing everyone, that doesn’t sound to me like science fiction.”
Rosie Campbell: “I just want some actual observational evidence showing AI causing human extinction”
Tagon: “Can you tell me how AI outsmarts all humans, but in a way I understand?”
Ash: If it is not a bulletproof, completely understandable by the average person, literally-kills-everyone scenario that I cannot find a hole in according to my own rubric, this AI doom stuff is all bullshit. /s
Often you get the classic ‘tell me exactly what moves Magnus Carlsen would use to beat me in chess, or I don’t think he would win’ style arguments, except people don’t realize that this is what they are asking.
This is an understandable but wrong question, as longtime readers here know.
You can see Geoffrey Hinton correctly trying to sidestep it here, in an interview with a clearly stunned BBC TV anchor. It is highly understandable that people ask. There is, in practice, no great way to respond.
Jasmine Sun calls for talk about ‘specific safety solutions’ rather than ‘vagueposting about extinction.’ This is not vagueposting. This is very explicit, clear, shouting from the rooftops that we are all going to die posting. What do you want, the precise mechanism?
I am of course all for specific safety solutions, but if we assume that particular prosaic actions can solve the problem, especially without everyone being sufficiently alarmed to consider actively costly interventions, then we are probably already dead.
Quite a lot of people, as always, fixate on ‘okay, but how exactly, physically, would AI do the particular things that caused everyone to die?’
They ask questions like this, often earnestly:
Frank Luntz: How would AI go about killing everyone on the planet?
David Hapunkt: Can you guys all please stop vague Posting and start explaining what the concrete risk is?
👶🏽🐐LaZr$ babyGOAT: Could someone explain to me how AI could “kill us all” without using a nuke or some type of biological means as an example? I’m not even saying I don’t believe it could, I just genuinely don’t understand how AI can literally kill us all. I mean I understand the social and economic impact to a degree but I’m talking about this doomsday scenario people keep warning us about but can’t quite seem to be specific about.
The problem, as Rob Bensinger puts it, is that if you need to do it without ‘sci-fi’ elements and with the rule that fiction has to make sense and the humans can’t make mistakes, then this is a deeply silly exercise. Not that the AI wouldn’t win anyway, but it’s a fundamentally wrong exercise that in 2025 would have excluded talking about versions of many of the actual events of 2026.
Here is Coxon’s answer in Wired, which I like a lot:
Maxwell Zeff: Can you draw a line for me between the alignment problem, which I think the Hugging Face incident is an example of, and something you said in your [Twitter] post, which is that “the people building AI earnestly believe that it could kill us all by the end of the decade.” I don’t think everyone understands how those are connected.
Jacob Coxon: The main obstacle to understanding this is that it sounds like science fiction. But it’s kind of important that everyone who writes science fiction about AI comes to the conclusion that there’s a big risk that a much smarter thing can kind of take over. We’ve got this as a trope, but there’s an obvious grain of truth to it.
Imagine you versus a monkey. AI has the same sort of difference in intelligence to a human as we do to a monkey, which I think is quite an extreme intellect difference.
And now, imagine that we have to control the behavior of this vastly smarter thing, which is the problem of alignment—ensuring that it does exactly what we want. It’s pretty difficult for a monkey to control a human, just by a kind of simple analogy. We have to be very careful that we get the control problem exactly right.
When we say human extinction, it’s because, for something that intelligent, it really will be quite straightforward for it to kill everyone. Imagine the AI decides it doesn’t want to be turned off, which I think is quite a natural thing for an AI not to want, right? For whatever reason, it decides it doesn’t want to end. And it realizes the human is gonna turn it off tomorrow. So how does it stop the human turning it off tomorrow? Maybe it’s got some clever way, but if it’s a sufficiently smart thing, it could just, you know, wipe out humanity so it doesn’t get turned off.
Most of the time, this is The Way.
Ideally, you get to answer with ‘recursive self-improvement,’ after which AIs that can outcompete and outthink us take control and we become irrelevant and then dead, maybe there are robots or whatever involved I guess if that helps, but the moment you say anything about robots or plagues or anything else physical you’re probably cooked, you mentioned specific things.
Another simple explanation, for the ‘how do the AIs take control’ part: People will find that things work better when you put AIs in charge of them, so AIs get put in charge of or gain access to essentially everything. They are handed control over all the resources. I mean, sure, it would happen without that, but that’s the simple way. Then, once AIs are in control of everything, and either cooperating with each other or competing (either way), they will automate their supply chains because that is more efficient. Then the world gets used for things AIs care about, and we will not be worth the costs of keeping us alive and dealing with us. So we die.
Those are kind of the basic, ‘good’ scenarios where something worse, or an earlier failure, does not happen to us first. In reality, expect earlier failure.
Typically, if you answer with specific things or an actual list of events, a skeptic will look for a reason to dismiss any particular physical causal chain as implausible, or simply find the whole thing silly because any given story will sound silly if you want it to sound silly, or as in this exchange between Kelsey Piper and Andy Semeza you get caught up in minutiae, and people try to apply the principle that the humans could beat this particular approach and win if the humans would just (which as we all know they won’t, humans never just), or even just say things like PoliMath’s ‘humans are good at surviving things,’ because the AIs would never then take any additional moves in such a game if you survived stage one.
When in reality obviously none of those details actually matter. Once the smarter thing that keeps getting more capable and doesn’t care so much about you is in charge, you are cooked. You might not be cooked today. You might not be cooked tomorrow, if you are lucky. You are still, once it matters whether or not you are cooked, or it would take effort to not cook you, about to be highly cooked.
There is no great simple answer that works consistently in practice. Many people have spent quite a lot of time trying to figure out how to best explain. There are books, and ** there are scenarios**.
Yes, the versions where all humans are literally dead by 2030 will require some sort of science fiction thing to happen, on top of all the science fiction things that have already happened. But does the exact date of the strictly final human dying matter?
One central problem is that different explanations work on different people, who have 50+ distinct particular objections and fixations, and when you are making a prediction or writing fiction your story ‘has to make sense’ and you cannot have people act half as stupid as people typically act, whereas reality has no such requirements.
Ultimately the physical details do not matter. AIs will compete with each other for resources, or they will cooperate with each other to take the resources, and otherwise try to achieve some set of goals. Because we will give them goals explicitly and on purpose, and also they will happen to have various goals. The best ways to accomplish those goals will not involve humans. So the humans will cease to have the resources required to sustain themselves. Then the humans will die. That’s it.
I realize that answer is unsatisfying and unconvincing to many, but I have too much else to cover to worry about that right now.
Best Start Believing In Science Fiction Stories Because You Are In One
You did notice that you live, today, in a science fiction world, right? There are swarms of agents hacking websites and solving Millennium problems. Stop pretending.
If you think all this talk is science fiction, think about what you would have said in 2023, or even 2025, let alone 2016, if I had described the events of 2026. You would absolutely have called it ‘science fiction.’ I mean, okay, sure, we can explain how you would die even without AI ever doing anything it can’t already do, only by doing it all sufficiently faster and better and cheaper, if that’s what you really want, like I need to get every move an AI makes past some sort of FDA-style approval process that always says no, but I don’t see the point.
The head of YC, Garry Tan, illustrates how crazy this mindset now looks, as he expresses views well summarized by Ball and Eth:
Dean W. Ball: “I don’t really care about science fiction… We need to actually talk about… what’s actually happening with the agent swarms” is the most perfect encapsulation of the vibes of Q3 2026 I have seen.
Daniel Eth (AI Safety): “AI extinction risk is simply a speculative concern that’s meant to distract us from more immediate risks like AI loss of control” was not even a take I had considered.
And here is the original statement from Tan:
Garry Tan: We should be talking less about this Jacob Coxon guy, and talking a lot more about — what is actually happening with Hugging Face? Are agent swarms going to take over infrastructure en masse? And then, what are we actually doing about that?
I don’t want to hear about some guy who worked for Anthropic for 2 months.
Tan has his facts wrong here. It was actually four months, as Coxon started in May, after years seeing the problems firsthand at OpenAI.
Garry Tan (resuming): There’s a coordinated effort to try to influence politicians to get a knee-jerk response out of them.
That’s a smokescreen. You shouldn’t be paying attention to that. We need to be paying attention to the actual things we can do to, for example, prevent agents swarms from taking over entire data centers. What’s our shutdown strategy? How do we ensure provenance? Where is this agent actually located? What software can we build? What cybersecurity defenses can we build today?
That’s the level of discourse I think we need, and we just don’t have that.
I don’t really care about science fiction. I saw Terminator 2, too. We’re not here to talk about that. We need to actually talk about what’s really happening with the servers, what’s actually happening with the agent swarms, and how do we actually prevent that?
QC: we really have to let go of the lazy idea that anything that sounds “sci-fi” is prima facie ridiculous and not worth considering as a possibility. it seems like the huggingface hack + navier-stokes has made this clearer recently, but nearly everything about AI capabilities here in september 2026 would have sounded like sci-fi to nearly everyone in any previous year, let alone the years before that (and you can check this by checking what expert forecasts looked like from 2025 or earlier, or checking what the reaction was to AI 2027 when it came out, or literally looking at old sci-fi with AI in it and noting which things the fictional AI do are also things that modern AI can now do)
but this pov is also missing the history of where sci-fi even came from as a genre. the progenitors of the golden age of science fiction (heinlein, campbell, asimov, etc) lived through an age of wonders: radio, TV, airplanes, computers, nukes, satellites, the moon landing. they were not just making shit up for fun! they wrote about robots and starships in part because they were surrounded by evidence that technological society was capable of inventing wonders and they expected further wonders! science fiction as a genre derives its strength from the empirical fact that we can and do invent powerful new technologies that completely alter the structure of human living, such as the technologies making it possible for you to read this right now. you live in a science fiction world and you always have
Literal Extinction Is Not Much Harder Than Loss of Control
The ‘well actually full literal extinction of every human is hard’ arguments are rather profoundly stupid when the opponent you face is superintelligence. If you need a non-sci-fi intuition pump: Was it hard for humans to drive so many other species to extinction, often while actively wanting not to do so?
As in:
roon (OpenAI): human extinction is not easy. none of the “respectable” ai risk factors people normally mention in the RSP / preparedness reports would get even close i think, like nuclear or pandemics. it has to be the grey goo. the parallel family of self-replicators
yes anything that replicates itself could in theory kill us but this sort of thing would leave a lot of time to respond —
Okay, ad arguendo, all you glacially slow idiot (relative to the AIs) fluid sacks out there, let’s say it doesn’t happen overnight. The AIs are in control of and have automated their supply chains. Go ahead and try to ‘respond.’ See how that goes.
Again, yes, if you reject all new tech moves as ‘science fiction’ then some of the humans do not literally die overnight. You will still be dead soon enough.
Conspiracytown Is Always Hiring
Dean W. Ball: “This has all the makings of a coordinated op funded by shadowy mega donors,” said people engaged in a coordinated op funded by shadowy mega donors.
Nate Silver: Debating process is usually a tell that you have a losing argument on the substance. And I don’t think this is one of the exceptions. Though I think Coxon is having a lot of impact, partly because the counterarguments have been so feeble.
Yep.
Parker Thayer (6m views) ended up at the forefront of the systematic effort to label Jacob Coxon part of some dark conspiracy, which is part of the yearslong efforts from certain corners to say that any warnings about risks from AI are ever and always part of that dark conspiracy.
I would prefer not to mention this, but it is making enough rounds that I will do so. We have to ensure that this ludicrous narrative does not take root, and cause these issues to become politically polarized. If you do not need to deal with such stupid claims, you can skip this section.
You see, before Coxon made news and tried to warn us, he went to The Wall Street Journal to ensure distribution of that news. This, says Thayer, is evidence of a conspiracy, rather than ‘how any normal whistleblower makes news, ever.’
Zac Hill: All the people being ~conspiratorial~ about the ‘well funded PR campaign’ (:P) behind the whistleblow should go talk to one (1) journalist, who rightly understands the playbook behind how to do something like this effectively.
Kelsey Piper: I would expect that he definitely reached out to journalists and to other ex-lab employees in order to get articles written about his decision to resign. This seems like a super obvious thing to do if you want to resign and whistleblow.
If I had something important I wanted to say and was not a journalist, I would reach out to lots of press outlets, talk to my friends and get their help drafting it, and then make a Twitter to tweet it (and probably ask some people to share it). Right??
For what it’s worth journalists love hearing from people who are like “I am quitting my job because they’re catastrophically irresponsible and what they’re doing should be banned”.
If I claimed to have some of the most important news ever, but I couldn’t be bothered with steps like ‘call a journalist’ and ‘ask some friends to retweet’ then you would quite rightfully wonder if I was taking any of this seriously.
It is laughable the straws being grasped. Look at all these people in AI Safety, Thayer says, who amplified Coxon’s message, like Nathan Calvin and Daniel Kokotajlo, and look at the (gasp) funding they have. Oh look, ‘it just so happened that’ 27 minutes after posting, someone from Coefficient Giving quoted Coxon. Are these not exactly the responses you would expect from those people, or myself, or many others, if they were online at the time?
Oh, and it was ‘punchy, quotable and professionally written.’ Yes, if you are going to burn your career to send a message, you are going to construct a good Tweet.
And ‘organic reach seems unlikely’ due to small following. This, as Theo says, is simply not how the modern Twitter algorithm works. Fully viral Tweets go viral because people engage with them. Going into the 100m+ view range has almost nothing to do with the original account’s reach. You literally cannot ‘manufacture’ 160 million views.
All that Coxon did, as per his interview, was arrange some early retweets from friends, to get this in front of the standard AI safety eyes. Twitter and the media did the rest, for a multitude of reasons I have described above. He explicitly confirms on Fox News he did not work with any third parties ‘in terms of this whistleblowing and coming out,’ which excludes only talking to the reporter at The Wall Street Journal, and consulting several friends.
Aella affirms that all the ‘usual suspects’ who would have loved for this to go viral, and indeed loved when it did go viral, were all extremely surprised when it went viral. None of us knew it was coming. I also affirm this. I had no idea who he was, until I saw the post, which of course I retweeted once I saw what it was.
Here we see Adam Townsend scroll through Tweet engagements and it looks exactly like you would expect from high levels of organic engagement, especially from people worried about AI killing everyone. I know what ‘coordinated news stories’ look like, and this is not it. The evidence for ‘this was coordinated’ was ‘no one agrees with me.’
The entire question here, and the suggestion that it matters, is utter madness. Coxon did not consult anyone other than a few friends, but why would it hurt your credibility to try and get others to amplify your message? Is that not what you would do if you had a super important message?
The final evidence Thayer presents is ‘basically every major Democrat politician and candidate has suddenly glommed on to this post,’ again evidence that this must be a conspiracy because it worked. Why would politicians latch onto a highly legible warning that was going heavily viral, about heavily distrusted companies? It must be a conspiracy. More than that, it must be a conspiracy by investors in Anthropic, as Thayer points out these dark links to multiple Anthropic investors.
Zac Hill: I don’t understand what you’re attempting to imply really. I mean obviously the guy is going to ensure his tweet gains reach before tweeting, but the rest of this is just “interested people are interested”, isn’t it?
Then there’s Jordan Schachtel (1.2m views) who says this has ‘all the signs of a highly coordinated op through doomer mega donors and the corporate media’ where those signs are… a Tweet that went viral, that media paid attention, and that he wasn’t at Anthropic for that long? Thus, mega donors.
When it turned out he got the timing wrong there, he retreated to saying it was two months, whereas actually Coxon worked at Anthropic for four months, two short of the equity cliff, after he previously quit OpenAI over safety concerns.
Which is exactly how you would behave if you worked at OpenAI, thought Anthropic was the responsible one, moved there and (as he says in the Wired interview) did not actively see corners being cut yet but found out that no, Anthropic is better but there is no responsible one.
Maxwell Zeff: [Coxon] claims that, in his experience, Anthropic operates more responsibly than OpenAI, but he expects both companies could cut corners in the future if nothing is done to slow their race for dominance.
Jacob Coxon (in that interview): I think Anthropic is far and away the most responsible player in the space. Having worked at both OpenAI and Anthropic, there is a night-and-day difference in the extent to which they’re taking the situation seriously.
… To give an example, executives at OpenAI won’t give you their exact pictures for what the world will look like. They’ll never say, like, “This is exactly why we’re doing this, and this is the way the world will look.” They won’t give concrete predictions. Anthropic will have executives and leadership making very clear predictions and discussing, like, details of company strategy with the whole company.
And part of the reason they can do this is because everyone at Anthropic is treating this like we’re on war footing. They don’t leak anything—nothing ever leaks from Anthropic [Editor’s note: Some stuff has leaked]. OpenAI has leaks every day.
Andrew Egger: “Dude works … at OpenAI for years, despairs of OpenAI’s approach to safety, jumps to the company that’s supposed to be all about safety, discovers things aren’t much better there, gets blackpilled, quits, and blows the whistle” seems like an intelligible sequence!
I’ve seen a lot of conspiracytown accusations in my day, from all sides of the political spectrum. This is Q-Anon level grasping at straws.
Elon Musk: I think the groundwork for this psy op (for lack of a better term) has been prepared for a long time. This was just the match that lit the fire.
The groundwork is called ‘actual reality and real risk.’ As agreed to by Elon Musk. Elon Musk has repeatedly said that we will definitely lose control of AI, and this is why he has to be the one to build it first. No, seriously.
He also himself as put a 10%-20% chance of human extinction.
The whole thing is super frustrating, on so many levels, although Derek is overreacting, this in no way eclipsed the real story:
Derek Thompson: The insidery media-studies question of whether Jacob Coxon worked with other AI safety folks to draw attention to his resignation is, bafflingly, eclipsing the actually-existentially-fundamental question of whether powerful AI might be dangerous.
I understand why people with no opinions on, or interest in, AI safety would be drawn to focus on the more gossipy aspect of the story. But (a) whistleblowing with media contacts is normal and fine; (b) dude says he didn’t do it; (c) it doesn’t matter at all if he did!
If, eg, RSI is imminent and bad, it doesn’t matter if Jacob worked with 0, 10, or 100 third parties to get his message out; he’s still in the right.
If AI safety fears are a stupid neurotic moral panic, it also doesn’t matter if Jacob texted 0 or 10,000 people about his plans to resign; he still did something dumb and histrionic. Elon Musk cofounded OpenAI in part bc he was worried big tech would mess up AI safety. Dario et al confounded Anthropic in part bc they were concerned OpenAI was messing up AI safety.
Conservative commentators today: Wait is AI safety a secret ~group project~ that ~multiple parties~ have collaborated on in ~private text messages~ and for which they have contributed ~money~?
Fucking obviously yes?! That’s literally the entire history of America’s frontier labs?
Ultimately it is a combination of people looking to do hit jobs against any concern for safety, because that’s 2026 for you, and also Derek Thompson’s theory here, that this is simply the sense that everything, everywhere, all at once is always an op:
Derek Thompson: Something I’ve learned about being too online in the last few years is that some people really truly cannot believe that other people might come to opinions about the world that are different from their own in a way that is natural and fair. Everything has to be an op.
There are all sorts of people—mostly conservative but even tech critics like McNamee—who insist that this AI safety moment is a psyop. Have people’s feelings about AI safety been shaped by a complex set of factors such as longstanding AI fears, worsening data center politics, the Hugging Face attack and its fallout, and now the Coxon resignation? No, no, no. They’re victims of a psyop.
It reminds me that I’ve had a distressingly high number of conversations with leftists who believe that the only reason America isn’t a socialist paradise is that Fox News and other conservative media do something special and evil and conspiratorial to brainwash half the country. Look, I do not like Fox News. But I think Fox News is successful because it says what many people have already been thinking. It shapes public opinion at the margins, but mostly it reflects biases back to viewers rather than invent them out of whole cloth.
There is a powerful, powerful urge to believe that the only reason the world doesn’t agree with you about literally every little thing is that small handful of media-adjacent elites are constantly doing “ops” on the undiscerning public. I don’t know, man. People’s attitudes are different from yours and mine, because … their lives are different. They accumulate different biases, different expectations, different values, different frameworks. Difference is normal, not an op.
There are definitely ops running around, but the distressing truth is that mostly no one is competent to run all but the simplest ops, and mostly life is improvised chaos.
Now You See It
To those finally being struck by how stupid are the reasons of those dismissing AI risks, I, Eliezer Yudkowsky, Nate Soares and so many others say: Welcome.
To our personal hell, to be clear. Still: Welcome!