Dario Amodei, Anthropic’s chief executive, on his personal website: But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.
My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.
My second concern is the OpenAI-Hugging Face incident (OAI-HF), in which a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance…
I’m therefore proposing a three-step plan with the goal of pacing the frontier: building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas. To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third-party evaluators to confirm this.
I have frequently said that the preeminent safety concern about artificial intelligence systems is content moderation: people doing nefarious things with the models in a way that is not aligned with human interests. Developing biological weapons, planning terrorist attacks, generating nonconsensual sexual imagery, and more. I am not recanting that belief; rather, I am amending it. Unlike many others who share this view, I do find the OpenAI-Hugging Face incident deeply alarming, and I’m glad it’s serving as a wake-up call to the industry. But to add to Amodei’s alarm, I think this incident and others like it are also fundamentally content moderation problems. After all, what is a content moderation failure if not for a system — whether it be automated, human-managed, or some combination of the two — acting against human interests?
Large language models are not sentient beings with their own thoughts, emotions, beliefs, and motives. The OpenAI model that hacked Hugging Face was not acting nefariously or in its own self-interest — it was acting in the interest of the task it was given, based on the instructions it was provided. The model was trained and designed poorly after OpenAI engineers — as reported numerous times in the media — forwent safety and alignment research to advance pure model capability. The system was designed execrably, and humans are to blame for that misalignment. Even if models in the future are created purely through recursive self-improvement, and those models become misaligned, that too is a human engineering failure. These models are our products; they do not have their own motivations. They act “autonomously” only via the instructions and pre-training and reinforcement learning we gave them.
If we think of the models as autonomous beings that are not products of our own work — that are not reflections of our own motivations and priorities — then yes, these incidents seem chilling. If we neglect the fact that humans made the models that went rogue, then yes, a kill switch to end AI development, like some politicians are suggesting, seems in order. But this is so detached from reality — we, the humans, are the ones who must be regulated. The technology companies abandoning safety and alignment research must face criminal responsibility for their models’ actions. If someone programs a computer virus and unleashes it on some network, we don’t write laws to develop a kill switch for viruses. We don’t write laws banning programs that delete files or log keystrokes. We write laws that say writing computer viruses and infecting others’ computers with them is a felony. This brings me to my final argument: The tech industry cannot possibly be trusted with regulation, because when profit clashes with ethics, corporations usually choose profit. While it’s moderately relieving that the industry’s two largest players have miraculously agreed on this one point — “pacing the frontier” — it’s unlikely that pattern will continue. Humanity’s safety and security should not rely on a few men in San Francisco having it in their hearts to do the right thing, and neither should it rely on a city of mentally deranged octogenarians to keep up with modern technology. There must be a better way.