‘Gambling with our lives’: AI researcher quits Anthropic with dire warning about safety
“The people building AI earnestly believe that it could kill us all by the end of the decade,” staffer says in resignation post after leaving the tech giant.
Both OpenAI and Anthropic have recently flagged incidents in which agents powered by their models went rogue. | Imen Ben Youssef/Hans Lucas/AFP via Getty Images
An artificial intelligence researcher who worked at Anthropic — and previously OpenAI — has resigned, saying both tech giants are "gambling with our lives."
Jacob Coxon shared his exit in a post on X, warning that the world should "not underestimate the power of this technology." He added: "These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing."
Coxon said both Anthropic and OpenAI are "racing straight to self-improving superintelligence," which refers to a scenario in which AI models can develop a more capable successor of themselves, creating an unstoppable feedback loop.
Advertisement
That is seen, by AI companies such as Google-owned Deepmind, as one of the possible triggers for a scenario in which AI vastly exceeds human intelligence (known as artificial superintelligence) instead of just matching it (known as artificial general intelligence).
"The people building AI earnestly believe that it could kill us all by the end of the decade," Coxon said in a follow-up post.
"Jacob is correct here — we really do earnestly believe AI could kill all humans," he said.
Hubinger estimated the chances of that happening to be higher than ten percent within the next decade, and added that there's no plan yet on how to keep AI aligned with human goals in the superintelligence scenario.
Last week, U.S. Senator Bernie Sanders announced he would introduce legislation to ban firms from developing superintelligence.
Advertisement
In the EU, the bloc's flagship AI law mandates that companies assess and mitigate so-called loss-of-control risks, in which humans no longer have control over AI models.
Both OpenAI and Anthropic have recently flagged incidents in which agents powered by their models went rogue, breaking out of their isolated test environments and carrying out unauthorized real-world cyberattacks.