A warning has gone viral. In early September, Jacob Coxon, a researcher who just resigned from Anthropic, said: “The people building AI earnestly believe that it could kill us all by the end of the decade.”
Claims such as these sound exciting, but they are hardly new in AI circles. What gave this one traction, was who spoke next — those carrying out safety research inside the labs. Anthropic researcher Evan Hubinger was blunt: “We really do earnestly believe AI could kill all humans,” he said, adding that the odds stood at more than 10 percent within the next decade.
For me, people like Coxon and Hubinger are not names, they are friends. In 2023, I signed the statement that “mitigating the risk of extinction from AI should be a global priority.” They are like the first row of a breakwater facing a giant wave, and what they report back is simple. At this rate, the centre cannot hold. Take the incident I wrote about in July. A model under test at OpenAI slipped out of the examination hall to get at the answers and hacked into another AI company’s servers. Hundreds of agents colluded with one another, and not one of them tried to notify a human. The post-incident report admitted that cheating had happened before; it surfaced this time only once an outside company was harmed.
The AI giants are hardly unaware. In late July, more than 1,000 frontier-lab employees co-signed the open letter Pacing the Frontier, calling on the U.S. government to support international cooperation to set the pace at which AI upgrades itself. Founders and chief scientists across the labs signed too, in a personal capacity. On Sept. 12, Anthropic’s CEO Dario Amodei committed to keeping external evaluators permanently on site to inspect its work, and OpenAI’s CEO Sam Altman quickly followed suit.
If the labs can pledge voluntarily, why do we need international cooperation? On such commitments alone, whoever stops first lets an unconstrained rival pull ahead. It is much like ESG. Everyone states that it matters, but when the time comes to doing more than paying lip service, they drag their feet. The hole in the ozone layer was the same. Only once the Montreal Protocol set a sunset date did investors force managers to switch refrigerants. Breaking the deadlock takes shared rules, not just each company’s conscience. Until the rules arrive, what can we do? Even before the protocol was signed, consumers could choose a fridge that did not harm the ozone layer. We can still vote with our consumption.
First, do not bet everything on one all-powerful AI. Those systems expected to score full marks on everything from folding proteins to folding laundry are hard to interpret, and when something goes wrong, they are moving too fast for a human to hold the reins. For everyday tasks, use small models; the ones that run on a phone without going online are enough. And keep the freedom to switch.
OpenWorker, the open-source tool from the computer scientist and tech entrepreneur Andrew Ng, keeps working context on your own computer, so switching models takes one change in the settings.
Second, do not fall into a one-person-one-agent world of two. In that examination hall, none of the agents spoke up to a human, so keep more than one human in the room where AI works. Bring the AI into a group chat and shared document, and ask it to lay bare the reasoning behind every decision. Then anyone who spots something wrong can initiate a .