{"slug": "the-ai-safety-vibe-shift", "title": "The AI  safety vibe shift", "summary": "Former Anthropic pretraining researcher Jacob Coxon resigned and posted on X that OpenAI and Anthropic are \"racing straight to self-improving superintelligence and gambling with our lives,\" a post that drew 159 million views. Anthropic alignment science lead Evan Hubinger amplified the claim, writing that he personally puts the risk of AI killing all humans at more than 10% within the next decade and that \"we do not yet have a plan to solve alignment for superintelligence,\" adding 41 million views. The episode follows OpenAI chief scientist Jakub Pachocki's blog post \"An Alien Mind\" warning about recursive self-improvement and the world's lack of preparation.", "body_md": "[AI Safety](https://www.platformer.news/tag/ai-safety/)\n\n# The AI safety vibe shift\n\nOnce a fringe obsession of Bay Area rationalists, existential risk is suddenly all anyone is talking about\n\n*This is a column about AI. My fiancé works at Anthropic. See my full ethics disclosure* *here**.*\n\nOn Monday, after a weekend’s worth of concerned blog posts from OpenAI, I wrote about [__why we ought to take those warnings__](https://www.platformer.news/openai-astra-warning-alignment-monitoring-pachocki/) seriously. For a host of reasons, I wrote, AI companies make for flawed messengers on this subject: they can reasonably be accused of marketing, of blame-shifting, of regulatory capture, and more. And so I can understand why, despite a rising drumbeat of ominous talk over the past several years, the mainstream has largely written off existential risk as a fantasy.\n\nMoments before I sent out that edition, though, a recently departed Anthropic researcher published an X post that brought the subject to the center of conversation. “I resigned from Anthropic today,” wrote a young researcher named Jacob Coxon, in a tweet that generated 159 million views. “I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”\n\nCoxon’s post picked up on the same themes as recent comments from OpenAI — including from chief scientist Jakub Pachocki, whose blog post “[__An Alien Mind__](https://openai.com/index/an-alien-mind/?ref=platformer.news)” also warned (in less lacerating terms) about the risks posed by recursive self-improvement and the world’s disturbing lack of preparation for the consequences. \n\nIt tells you something about the bubble I inhabit that, for those reasons, Coxon’s post initially did not make much of an impression on me. He is hardly the first Anthropic employee to resign for safety-related reasons; in February there was a minor stir after a fellow safety researcher [__quit to pursue__](https://www.bbc.com/news/articles/c62dlvdq3e3o?ref=platformer.news) a poetry degree after becoming distressed about the pace of progress and Anthropic’s role in it. \n\nAnthropic is a company founded by people who believed their former colleagues at OpenAI paid insufficient attention to safety and that they had at best an outside shot of steering the world to a better outcome. More than three years ago, my Hard Fork co-host Kevin Roose profiled the company and called it “[__the white-hot center of AI doomerism__](https://www.nytimes.com/2023/07/11/technology/anthropic-ai-claude-chatbot.html?ref=platformer.news).” For Anthropic, that sense of doom attracts talent, gives them a sense of purpose, and (less often) ultimately drives them away.\n\nUntil recently, this was mostly a thing people made fun of the company for. But shortly after Coxon’s original post, it got a signal boost from a striking source — Evan Hubinger, who continues to lead alignment science at the company.\n\n“Jacob is correct here—we really do earnestly believe AI could kill all humans!” [__Hubinger wrote__](https://x.com/EvanHub/status/2097497037956891126?ref=platformer.news). “I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”\n\nThat was good enough for another 41 million views and another jolt to the discourse, despite the fact that (as Nitasha Tiku points out at [__the *Washington Post*__](https://www.washingtonpost.com/technology/2026/09/10/years-they-warned-ai-could-kill-all-humans-now-people-are-listening/?pwapi_token=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJyZWFzb24iOiJnaWZ0IiwibmJmIjoxNzg5MDEyODAwLCJpc3MiOiJzdWJzY3JpcHRpb25zIiwiZXhwIjoxNzkwMzk1MTk5LCJpYXQiOjE3ODkwMTI4MDAsImp0aSI6IjBkOTMwOTI3LTEzMWEtNGVlMC1iYzYyLWUwZDlmOWQzYzQ4OCIsInVybCI6Imh0dHBzOi8vd3d3Lndhc2hpbmd0b25wb3N0LmNvbS90ZWNobm9sb2d5LzIwMjYvMDkvMTAveWVhcnMtdGhleS13YXJuZWQtYWktY291bGQta2lsbC1hbGwtaHVtYW5zLW5vdy1wZW9wbGUtYXJlLWxpc3RlbmluZy8ifQ.Cr7buGMB5cztSS6bQG7OmkeAY0yJhfOg6kvY0kMj17c&ref=platformer.news)) Hubinger had posted warnings like this for years. (“My guess is that … when we put it in a situation where it thinks it can kill us, it just murders us,” he said in a 2022 talk.)\n\nOf course, back then, Anthropic wasn’t a $965 billion company heading toward what could be the biggest initial public offering of all time. Claude didn’t exist, and its models had not yet [__broken out of their sandboxes__](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals?ref=platformer.news) to compromise at least three separate organizations. (The company [__said Wednesday__](https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents?ref=platformer.news) that it had hired research organization METR to conduct an investigation of those incidents.)\n\nIn short, AI felt less serious then. But in the aftermath of [__the OpenAI attack on Hugging Face__](https://www.platformer.news/openai-huggingface-metr-report-slowdown/) and the general turn in public opinion against AI, the public seems increasingly ready to acknowledge the risks of building superhuman intelligence. Fortunately, some members of Congress do, too.\n\nEarlier this month, Sen. Bernie Sanders (I-Vt.) and Rep. Greg Casar (D-Texas) [__announced__](https://www.sanders.senate.gov/press-releases/news-sanders-casar-introduce-legislation-to-ban-artificial-superintelligence-and-temporarily-pause-advanced-ai-development/?ref=platformer.news)  the Ban Artificial Superintelligence Act. If signed into law, the act would ban companies from developing superintelligence and enforce a temporary pause in advanced AI development. It would also instruct the US government to seek international agreements that would prevent the development of superintelligence globally. \n\n“Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with potentially cataclysmic results,” Sanders said, accurately. “The leaders of the major AI companies publicly acknowledge that they do not fully understand the technology and that it is escaping their control. It is irresponsible for society to allow them to move forward and make these products even more advanced.”\n\nFor the moment, it is unclear whether Sanders’ legislation has much chance of making it onto President Trump’s desk — or whether this oligarch-friendly administration would sign it. But public sentiment around AI is changing rapidly, and it is only really moving in one direction.\n\nThe concern is bipartisan, too. On Thursday, *Axios* reported that Sen. Josh Hawley — chair of the Senate Homeland Security & Governmental Affairs subcommittee on Disaster Management — had [__opened a probe__](https://www.axios.com/2026/09/10/openai-hugging-face-senate-investigation-hawley?ref=platformer.news) into the Hugging Face attack. Hawley reportedly called OpenAI’s handling of the situation “reckless.” “The American people deserve to know the details of what went on in the Hugging Face incident and other incidents of AI models going rogue,\" he wrote. (He also noted that \"just this week, three Anthropic researchers expressed publicly that there is a greater than 10% chance that AI could kill all human beings within the next decade.\")\n\nSen. Ted Cruz is also now reportedly [__working on__](https://www.washingtonpost.com/technology/2026/09/09/anthropic-researcher-resigns-warning-reckless-race-toward-superintelligence/?ref=platformer.news) legislation to address the catastrophic risks of AI. And there are also signs that the labs’ oft-stated preference for some sort of slowdown has teeth. In particular, safety advocates were cheered this week by news that Paul Christiano, an influential safety and alignment researcher and adviser to the US government, [__had joined__](https://techcrunch.com/2026/09/09/openai-adds-a-prominent-ai-doomer-to-its-board-of-directors/?ref=platformer.news) the OpenAI Foundation’s board and its safety and security committee. Among other things, that committee governs whether OpenAI’s for-profit arm can release new models; Christiano’s presence has raised hopes that it might slow some of those releases.\n\nNow, excited public conversation (and in particular, public conversation on X) is no substitute for actual regulation. Congress has a long history of introducing tech regulations that go nowhere — and of launching probes that go nowhere. And if Christiano’s best efforts at getting OpenAI to self-regulate fall short — well, he won’t be the first board member to have that experience.\n\nStill, I can’t help but be heartened by the vibe shift in AI safety. Last month, in the wake of Mark Zuckerberg’s glad-handing treatise on “personal superintelligence,” I argued that [__true superintelligence is a dragon__](https://www.platformer.news/zuckerberg-ai-manifesto-dragons/). And while I remain hopeful that it can be tamed, I believe the researchers who say we are nowhere close to being sure of it — and are quitting, in protest, jobs that would make them rich. The right people are now listening to them, and here’s hoping they take action.\n\n**Elsewhere in AI safety:** I can't decide whether it is deeply perfect or completely insane that Anthropic chose this day of all days reveal that this year it has already disrupted several plots by scientists that could have helped to [develop biological weapons](https://www.nytimes.com/2026/09/10/us/politics/anthropic-ai-biological-weapons.html?unlocked_article_code=1.AFE.vqLg.utU6-8AVkIUw&smid=nytcore-ios-share&ref=platformer.news), as part of a larger \"[threat intelligence report](https://www.anthropic.com/threat-intelligence-report-september-2026?ref=platformer.news)\" detailing efforts to use Claude for surveillance, cyberattacks, influence operations, and more.\n\n**On the podcast this week:** We are actually *off* this week in preparation for our final episode, but we have an excellent conversation between Ezra Klein and Jasmine Sun about data centers to tide you over.\n\n**Apple** **|** **Spotify** **|** **Stitcher** **|** **Amazon** **|** **Google** **|** **YouTube**\n\n## Following\n\n### Anthropic’s economists think we can *maybe* keep our jobs\n\n**What happened:**  **Anthropic** economists released a [__new interactive model__](https://www.anthropic.com/institute/econ-scenarios?ref=platformer.news) of AI-based economic disruption. The model takes in predictions about the trajectory of AI capabilities and outputs scenarios about how the economy and jobs will fare in 2030 as a result.\n\nThe researchers present a “moderate,” a “substantial,” and an “extreme” scenario.\n\nInterestingly, the team’s “extreme” scenario, which predicts GDP growth rising to a truly bonkers 15.4% per year in 2030, would lead to an 8.9% increase in unemployment from AI.\n\nMeanwhile, their “substantial” scenario, which relies on AI predictions based on a recent survey of 10,000 Americans, predicts a still-high 5.4% growth rate in 2030, but only an 0.7% increase in unemployment from AI.\n\nThe forecasts are substantially more conservative than the warnings CEO **Dario Amode** i has issued regarding AI-related unemployment. Last year, Amodei [__told *Axios*__](https://www.axios.com/2025/05/28/ai-jobs-white-collar-unemployment-anthropic?ref=platformer.news) that he worries about a scenario where “Cancer is cured, the economy grows at 10% a year, the budget is balanced — and 20% of people don't have jobs.”\n\n**Why we’re following:** It’s interesting to see Anthropic’s economists being (relatively) optimistic here. One reason why: they expect AI-related economic growth to increase demand and wages for manual labor, offsetting losses in computer-based work like software engineering.\n\nOne of the authors' main conclusions, though, is that, given a range of all seemingly reasonable parameters, “the range of outcomes is wide.” So we could see a scenario with substantial growth and modestly higher unemployment, or crazy growth and crazy unemployment.\n\n**What people are saying:** \n\n**LSE** economics professor **Ben Moll** and **Google DeepMind**’s director of AGI economics **Alex Imas** [__wrote an essay on **Substack**__](https://aleximas.substack.com/p/will-ai-soon-lead-to-double-digit) dismissing the predictions of double-digit GDP growth that Amodei and others [have made](https://www.youtube.com/watch?v=K7F6ohcBJus&ref=platformer.news). The economists think that such extreme forecasts would assume too much: Double-digit growth rests on several shaky assumptions, they write, including“machines to do most of the economy’s work by 2035, people to keep spending on whatever gets automated, buyers and investors who absorb the new output, [and] AI that never destroys value or slows its own deployment.”\n\n“Each of these assumptions may fail in the real world,” the authors wrote.\n\nThey added that “even Anthropic's own economic model only reaches double-digit growth in its \"extreme\" scenario.”\n\n—*Ella Markianos*\n\n### Those good posts\n\n*For more good posts every day,* *follow Casey’s Instagram stories**.*\n\n([Link](https://www.threads.com/@joy.scaglione/post/DdFz3gXjU4Q?ref=platformer.news))\n\n([Link](https://www.threads.com/@zachbroussard/post/DdFlR55HfeJ?ref=platformer.news))\n\n([Link](https://www.threads.com/@okpants/post/DdDygOJkapo?ref=platformer.news))\n\n([Link](https://www.threads.com/@seattlefoodgeek/post/DdAwEqtm9-M?ref=platformer.news))\n\n### Talk to us\n\nSend us tips, comments, questions, and AI safety discourse: [casey@platformer.news](mailto:casey@platformer.news). Read [our ethics policy here](https://www.platformer.news/ethics/).", "url": "https://wpnews.pro/news/the-ai-safety-vibe-shift", "canonical_source": "https://www.platformer.news/ai-safety-vibe-shift-coxon-anthropic/", "published_at": "2026-09-11 01:04:25+00:00", "updated_at": "2026-09-11 01:23:16.807904+00:00", "lang": "en", "topics": ["ai-safety", "artificial-intelligence", "ai-policy"], "entities": ["Anthropic", "OpenAI", "Jacob Coxon", "Evan Hubinger", "Jakub Pachocki", "An Alien Mind", "Platformer", "Washington Post"], "alternates": {"html": "https://wpnews.pro/news/the-ai-safety-vibe-shift", "markdown": "https://wpnews.pro/news/the-ai-safety-vibe-shift.md", "text": "https://wpnews.pro/news/the-ai-safety-vibe-shift.txt", "jsonld": "https://wpnews.pro/news/the-ai-safety-vibe-shift.jsonld"}}