cd /news/ai-safety/the-ai-safety-vibe-shift · home topics ai-safety article
[ARTICLE · art-126405] src=platformer.news ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

The AI safety vibe shift

Former Anthropic pretraining researcher Jacob Coxon resigned and posted on X that OpenAI and Anthropic are "racing straight to self-improving superintelligence and gambling with our lives," a post that drew 159 million views. Anthropic alignment science lead Evan Hubinger amplified the claim, writing that he personally puts the risk of AI killing all humans at more than 10% within the next decade and that "we do not yet have a plan to solve alignment for superintelligence," adding 41 million views. The episode follows OpenAI chief scientist Jakub Pachocki's blog post "An Alien Mind" warning about recursive self-improvement and the world's lack of preparation.

by read9 min views2 publishedSep 11, 2026
The AI  safety vibe shift
Image: Platformer (auto-discovered)

AI Safety Once a fringe obsession of Bay Area rationalists, existential risk is suddenly all anyone is talking about

This is a column about AI. My fiancé works at Anthropic. See my full ethics disclosure here*.*

On Monday, after a weekend’s worth of concerned blog posts from OpenAI, I wrote about why we ought to take those warnings seriously. For a host of reasons, I wrote, AI companies make for flawed messengers on this subject: they can reasonably be accused of marketing, of blame-shifting, of regulatory capture, and more. And so I can understand why, despite a rising drumbeat of ominous talk over the past several years, the mainstream has largely written off existential risk as a fantasy.

Moments before I sent out that edition, though, a recently departed Anthropic researcher published an X post that brought the subject to the center of conversation. “I resigned from Anthropic today,” wrote a young researcher named Jacob Coxon, in a tweet that generated 159 million views. “I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”

Coxon’s post picked up on the same themes as recent comments from OpenAI — including from chief scientist Jakub Pachocki, whose blog post “An Alien Mind” also warned (in less lacerating terms) about the risks posed by recursive self-improvement and the world’s disturbing lack of preparation for the consequences.

It tells you something about the bubble I inhabit that, for those reasons, Coxon’s post initially did not make much of an impression on me. He is hardly the first Anthropic employee to resign for safety-related reasons; in February there was a minor stir after a fellow safety researcher quit to pursue a poetry degree after becoming distressed about the pace of progress and Anthropic’s role in it.

Anthropic is a company founded by people who believed their former colleagues at OpenAI paid insufficient attention to safety and that they had at best an outside shot of steering the world to a better outcome. More than three years ago, my Hard Fork co-host Kevin Roose profiled the company and called it “the white-hot center of AI doomerism.” For Anthropic, that sense of doom attracts talent, gives them a sense of purpose, and (less often) ultimately drives them away.

Until recently, this was mostly a thing people made fun of the company for. But shortly after Coxon’s original post, it got a signal boost from a striking source — Evan Hubinger, who continues to lead alignment science at the company.

“Jacob is correct here—we really do earnestly believe AI could kill all humans!” Hubinger wrote. “I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

That was good enough for another 41 million views and another jolt to the discourse, despite the fact that (as Nitasha Tiku points out at the Washington Post) Hubinger had posted warnings like this for years. (“My guess is that … when we put it in a situation where it thinks it can kill us, it just murders us,” he said in a 2022 talk.)

Of course, back then, Anthropic wasn’t a $965 billion company heading toward what could be the biggest initial public offering of all time. Claude didn’t exist, and its models had not yet broken out of their sandboxes to compromise at least three separate organizations. (The company said Wednesday that it had hired research organization METR to conduct an investigation of those incidents.)

In short, AI felt less serious then. But in the aftermath of the OpenAI attack on Hugging Face and the general turn in public opinion against AI, the public seems increasingly ready to acknowledge the risks of building superhuman intelligence. Fortunately, some members of Congress do, too.

Earlier this month, Sen. Bernie Sanders (I-Vt.) and Rep. Greg Casar (D-Texas) announced  the Ban Artificial Superintelligence Act. If signed into law, the act would ban companies from developing superintelligence and enforce a temporary in advanced AI development. It would also instruct the US government to seek international agreements that would prevent the development of superintelligence globally.

“Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with potentially cataclysmic results,” Sanders said, accurately. “The leaders of the major AI companies publicly acknowledge that they do not fully understand the technology and that it is escaping their control. It is irresponsible for society to allow them to move forward and make these products even more advanced.”

For the moment, it is unclear whether Sanders’ legislation has much chance of making it onto President Trump’s desk — or whether this oligarch-friendly administration would sign it. But public sentiment around AI is changing rapidly, and it is only really moving in one direction. The concern is bipartisan, too. On Thursday, Axios reported that Sen. Josh Hawley — chair of the Senate Homeland Security & Governmental Affairs subcommittee on Disaster Management — had opened a probe into the Hugging Face attack. Hawley reportedly called OpenAI’s handling of the situation “reckless.” “The American people deserve to know the details of what went on in the Hugging Face incident and other incidents of AI models going rogue," he wrote. (He also noted that "just this week, three Anthropic researchers expressed publicly that there is a greater than 10% chance that AI could kill all human beings within the next decade.")

Sen. Ted Cruz is also now reportedly working on legislation to address the catastrophic risks of AI. And there are also signs that the labs’ oft-stated preference for some sort of slowdown has teeth. In particular, safety advocates were cheered this week by news that Paul Christiano, an influential safety and alignment researcher and adviser to the US government, had joined the OpenAI Foundation’s board and its safety and security committee. Among other things, that committee governs whether OpenAI’s for-profit arm can release new models; Christiano’s presence has raised hopes that it might slow some of those releases.

Now, excited public conversation (and in particular, public conversation on X) is no substitute for actual regulation. Congress has a long history of introducing tech regulations that go nowhere — and of launching probes that go nowhere. And if Christiano’s best efforts at getting OpenAI to self-regulate fall short — well, he won’t be the first board member to have that experience.

Still, I can’t help but be heartened by the vibe shift in AI safety. Last month, in the wake of Mark Zuckerberg’s glad-handing treatise on “personal superintelligence,” I argued that true superintelligence is a dragon. And while I remain hopeful that it can be tamed, I believe the researchers who say we are nowhere close to being sure of it — and are quitting, in protest, jobs that would make them rich. The right people are now listening to them, and here’s hoping they take action.

Elsewhere in AI safety: I can't decide whether it is deeply perfect or completely insane that Anthropic chose this day of all days reveal that this year it has already disrupted several plots by scientists that could have helped to develop biological weapons, as part of a larger "threat intelligence report" detailing efforts to use Claude for surveillance, cyberattacks, influence operations, and more.

On the podcast this week: We are actually off this week in preparation for our final episode, but we have an excellent conversation between Ezra Klein and Jasmine Sun about data centers to tide you over.

Apple | Spotify | Stitcher | Amazon | Google | YouTube

Following #

Anthropic’s economists think we can maybe keep our jobs

What happened:  Anthropic economists released a new interactive model of AI-based economic disruption. The model takes in predictions about the trajectory of AI capabilities and outputs scenarios about how the economy and jobs will fare in 2030 as a result.

The researchers present a “moderate,” a “substantial,” and an “extreme” scenario.

Interestingly, the team’s “extreme” scenario, which predicts GDP growth rising to a truly bonkers 15.4% per year in 2030, would lead to an 8.9% increase in unemployment from AI.

Meanwhile, their “substantial” scenario, which relies on AI predictions based on a recent survey of 10,000 Americans, predicts a still-high 5.4% growth rate in 2030, but only an 0.7% increase in unemployment from AI.

The forecasts are substantially more conservative than the warnings CEO Dario Amode i has issued regarding AI-related unemployment. Last year, Amodei told Axios that he worries about a scenario where “Cancer is cured, the economy grows at 10% a year, the budget is balanced — and 20% of people don't have jobs.”

Why we’re following: It’s interesting to see Anthropic’s economists being (relatively) optimistic here. One reason why: they expect AI-related economic growth to increase demand and wages for manual labor, offsetting losses in computer-based work like software engineering.

One of the authors' main conclusions, though, is that, given a range of all seemingly reasonable parameters, “the range of outcomes is wide.” So we could see a scenario with substantial growth and modestly higher unemployment, or crazy growth and crazy unemployment.

What people are saying:

LSE economics professor Ben Moll and Google DeepMind’s director of AGI economics Alex Imas wrote an essay on Substack dismissing the predictions of double-digit GDP growth that Amodei and others have made. The economists think that such extreme forecasts would assume too much: Double-digit growth rests on several shaky assumptions, they write, including“machines to do most of the economy’s work by 2035, people to keep spending on whatever gets automated, buyers and investors who absorb the new output, [and] AI that never destroys value or slows its own deployment.”

“Each of these assumptions may fail in the real world,” the authors wrote.

They added that “even Anthropic's own economic model only reaches double-digit growth in its "extreme" scenario.”

Ella Markianos

Those good posts

For more good posts every day, follow Casey’s Instagram stories*.*

([Link](https://www.threads.com/@joy.scaglione/post/DdFz3gXjU4Q?ref=platformer.news))

([Link](https://www.threads.com/@zachbroussard/post/DdFlR55HfeJ?ref=platformer.news))

([Link](https://www.threads.com/@okpants/post/DdDygOJkapo?ref=platformer.news))

([Link](https://www.threads.com/@seattlefoodgeek/post/DdAwEqtm9-M?ref=platformer.news))

Talk to us

Send us tips, comments, questions, and AI safety discourse: casey@platformer.news. Read our ethics policy here.

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-ai-safety-vibe-s…] indexed:0 read:9min 2026-09-11 ·