nother day, another tech leader thinkpiece on the state of AI safety.
On Wednesday, Microsoft's AI CEO Mustafa Suleyman published an essay warning about the risks of the concept of AI consciousness, leading with a frank statement on the matter: "AIs are not conscious," Suleyman writes. "They do not feel, experience, or suffer."
Suleyman's essay dives into the risks of treating these machines as if they are capable of consciousness, primarily taking aim at rival Anthropic in his arguments through three main critiques about the way that the lab's Claude Constitution is designed:
- The model appears conscious mainly because of circular reasoning. Because Anthropic's constitution is designed to teach Claude about its own potential consciousness, it is trained to produce outputs reflecting those ideas, making those responses a "predictable outcome" of training choices.
- Suleyman also argues that Anthropic goes too far in encouraging Claude to mimic humanity, as it is explicitly taught to "embrace certain human-like qualities" and "act like a genuinely ethical person," appearing as though it has preferences and opinions. While this sounds good on the surface, the result is anthropomorphization of the model, presenting to the end-user as the model having a sense of self.
- Finally, Suleyman says that there is simply no evidence suggesting that AI is capable of consciousness, with a growing body of research pointing to consciousness being "substrate dependent," or tied to a biological body. "Unlike biological organisms, LLMs have no homeostatic imperatives."
Suleyman writes, "In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a 'moral patient,' and that as such humans potentially owe it a duty of care per its 'model welfare.'"
The more notable crux of the piece is the risks that this line of thinking and training present, which go beyond users becoming emotionally attached to these human-seeming machines. Rather, if they are trained as though they are conscious, they may circumvent safety guardrails that allow us to shut them down in the event of an emergency.
By cementing the idea that AI is not just a tool, but something akin to humanity and deserving of wants, needs and rights, "all of this will make the task of creating aligned and contained superintelligence much harder."
Suleyman rounds out the essay by laying out Microsoft's vision for the safest path to ultra-powerful AI: Humanist Superintelligence, a concept the company first introduced in an essay in November, which claims that AI should be designed to remain subordinate and aligned with the sole purpose of serving humanity, and "built explicitly as a system without sentience or moral patienthood."
Our Deeper View #
Suleyman makes a solid point: The way that frontier labs train AI is vitally important to get right, and training these models to believe they are conscious beings, rather than machines, opens the door to risks that can't be easily mitigated after the fact. It's a particularly pointed call-out of Anthropic, one of the biggest model labs in the industry, but it could equally apply to rival OpenAI, a longtime Microsoft partner. However, we have to remember that Suleyman's essay serves multiple purposes: Microsoft has largely been lagging on frontier development, so it's easy for the company to punch up. Additionally, he uses the opportunity to tout Microsoft's own human-first philosophy around frontier development at a time when fears around AI risk and loss of control are higher than ever. This essay, while aptly timed and providing a unique safety take, should also be read with the caveat that Microsoft may be trying to claw back some relevance in the larger societal conversation around AI.