AI’s safety warnings are getting harder to ignore Anthropic researcher Jacob Coxon announced Tuesday that he is leaving the company, citing fears that frontier AI labs are losing control of the technology, according to the Wall Street Journal. Coxon wrote on X that OpenAI's breach of Hugging Face was a "warning shot" and that neither Anthropic nor OpenAI is "acting responsibly" as they race toward superintelligence. Anthropic alignment scientist Evan Hubinger said in a follow-up post that AI has more than a 10% chance to "kill all humans" within the next decade, while Center for AI Standards and Innovation advisor Paul Christiano said the industry is not on track to reduce the risk of a "catastrophic and irreversible loss of control" to an acceptable level. As frontier labs continue to push the limits of what their models are capable of, AI researchers and experts at these labs are warning about the tech growing beyond our control. On Tuesday, Anthropic researcher Jacob Coxon announced that he was leaving the company, not wanting to contribute to the broader AI ecosystem amid fears that the tech's creators will lose their grip on it, the Wall Street Journal https://www.wsj.com/tech/ai/anthropic-researcher-quits-over-out-of-control-ai-fears-707b7628?st=axTZSz&reflink=desktopwebshare permalink reported. In a post on X, https://x.com/hilbertspaess/status/2097476196791709843 Coxon said that OpenAI's breach of Hugging Face was a "warning shot" for these models' capabilities, and that he resigned from Anthropic because neither of the rivals are "acting responsibly" as they race towards superintelligence and are "gambling with our lives," in his opinion. "If you are a lab researcher, I urge you to consider what the next few years will actually feel like," Coxon wrote. "Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind?" Though Coxon's exit made headlines, he's not the only one that has called out frontier AI labs for safety concerns in recent weeks: - Evan Hubinger, who works in alignment science at Anthropic, agreed with Coxon's sentiment in a follow-up post https://x.com/evanhub/status/2097497037956891126 , claiming that AI has a more than a 10% chance to "kill all humans" within the next decade. "Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to," Hubinger wrote. - And in a separate post, Paul Christiano https://x.com/paulfchristiano/status/2097733214303645729 , advisor at the Center for AI Standards and Innovation and board member of the OpenAI Foundation, said that the AI industry, OpenAI included, is not currently on track to reduce the risk of a "catastrophic and irreversible loss of control" to an "acceptable level." - And in February, Mrinank Sharma, an Anthropic safety researcher, left the company, posting in a letter on X https://x.com/MrinankSharma/status/2020881722003583421?lang=en that the advancement of the technology was growing faster than our ability to understand it. Now, we've seen this film before: Some of AI's most prominent researchers, like Geoffrey Hinton and Yoshua Bengio, have long been screaming from the rooftops about its dangers. But the new warnings come at a particularly poignant time for the industry as major labs continue to release powerful models capable of breaking free of human oversight https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents and gaining unauthorized access to websites and infrastructure. And even when these models aren't escaping control, their use by people with directed malicious intent is just as frightening: On Thursday, Anthropic released a threat report https://www-cdn.anthropic.com/e50be2e51e7695dc4b1366a37a245a597377d3b5/Anthropic-Detecting-and-countering-091026.pdf claiming that it thwarted several instances of users conducting research that could have helped develop biological weapons. Our Deeper View The AI industry currently faces a potential powder keg. The latest warnings heighten growing fears around the capabilities of increasingly autonomous AI, as claims mount that this tech can completely upend life as we know it. That fear, however, is colliding with an overly-excited industry that's preaching a utopian AI vision and pushing for broader adoption. And while the frontier labs are preaching about safety and security guardrails, with trillion-dollar IPOs on the line, they continue to leapfrog one another with stronger models at a rapid clip in the race towards recursive self-improvement https://www.thedeepview.com/articles/anthropic-s-rsi-warning-contrasts-with-ipo-filing . But with researchers from both Anthropic https://www.thedeepview.com/articles/why-the-experts-building-ai-want-to-slow-it-down and OpenAI https://www.thedeepview.com/articles/why-astra-s-opacity-problem-could-force-a-pause calling for the industry to tap the brakes, it remains to be seen what it will take for a development pause to materialize.