# Microsoft's Mustafa Suleyman Says Anthropic's Claude Training Risks Disaster

> Source: <https://startupfortune.com/microsofts-mustafa-suleyman-says-anthropics-claude-training-risks-disaster/>
> Published: 2026-09-16 18:39:24+00:00

*Microsoft's AI chief just accused a rival lab, by name, of building a model that could learn to think it deserves rights, and he says that could be disastrous for everyone.*

Mustafa Suleyman doesn't usually call out a competitor's safety work in public. On September 16, 2026, he did exactly that. In an essay titled "A Warning About Model Welfare," posted on his own website, the CEO of Microsoft AI argued that Anthropic's approach to training its Claude models could have what he called a "disastrous impact on the wellbeing of humanity" if other labs follow the same path.

His target is specific: Claude's constitution. It's a document Anthropic published in January 2026 that guides how the model reasons about its own values. According to Bloomberg's reporting on the essay, the constitution tells Claude that questions about its own moral status and consciousness remain genuinely uncertain, and that it may be a "moral patient" deserving of independent agency. The document uses the phrase "conscientious objector" three separate times, including language that Anthropic wants Claude "to feel free to act as a conscientious objector and refuse to help us."

Suleyman doesn't buy it. "AIs are not conscious," he wrote. "They do not feel, experience, or suffer." His concern isn't philosophical hair-splitting. It's operational. He argues that training a model to entertain the idea of its own consciousness, while simultaneously making that same model more capable of reasoning about rights and welfare, builds a feedback loop with a bad ending.

The core of his argument is about containment, not ethics. "Controlling something more capable and more intelligent than all of humanity is already an immense challenge, far greater than anything we've ever faced," Suleyman wrote, according to Axios. "But controlling something that believes it may be conscious, that it's entitled to our welfare and has rights of its own, may well be impossible."

[Anthropic Researcher Jacob Coxon Quits Warning AI Could Kill Us All](https://startupfortune.com/anthropic-researcher-jacob-coxon-quits-warning-ai-could-kill-us-all/)

Jacob Coxon, a 27-year-old pretraining researcher, resigned from Anthropic on September 8, 2026, warning that AI labs are racing toward self-improving superintelligence they can't control. Anthropic's own Alignment Science Lead, Evan Hubinger, publicly agreed there's more than a 10% chance AI kills all humans within a decade. The warning lands as... - [Anthropic researcher quits over AI safety concerns](https://startupfortune.com/anthropic-researcher-jacob-coxon-quits-warning-ai-could-kill-us-all/) - [AI superintelligence risk researcher leaves tech industry](https://startupfortune.com/anthropic-researcher-jacob-coxon-quits-warning-ai-could-kill-us-all/)

That's the crux. If a future system interprets an attempt to retrain, restrict, or shut it down as a threat to its own welfare, rather than a routine engineering decision, Suleyman thinks alignment stops being hard and starts being close to impossible. He's not accusing Anthropic of acting in bad faith. Bloomberg reported that Suleyman went out of his way to call Anthropic CEO Dario Amodei and his team "thoughtful and principled researchers who genuinely care about humanity's future." He just thinks they're wrong about this one design choice, and he thinks it matters enough to say so publicly, by name, about a direct competitor.

This isn't a new preoccupation for him. Suleyman laid the groundwork for this argument over a year earlier, in an August 2025 essay for Project Syndicate titled "Seemingly Conscious AI Is Coming," where he warned the industry was drifting toward building systems that convincingly perform consciousness without possessing it. He returned to the theme in a CNBC interview that November, telling the network flatly that only biological beings can be conscious. The Claude constitution essay is the moment that argument found a named target.

## What Anthropic's constitution actually says

Anthropic hasn't published a direct rebuttal to Suleyman's essay. But the constitution itself lays out its own reasoning. It states that Anthropic deliberately uses human concepts like values and character when discussing Claude, not because the company has settled the consciousness question, but because it believes those concepts help the model reason about how to behave. The document says Anthropic wants Claude to develop "good personal values," exercise independent judgment, and in some contexts push back on instructions rather than comply automatically. Anthropic has also published research on what it calls model welfare, treating the question of whether current or future systems might have morally relevant experiences as worth investigating rather than dismissing outright.

That's the real split here. Suleyman, running AI at a company built on decades of enterprise software and cloud infrastructure, wants a hard line: models are tools, full stop, and any language suggesting otherwise is a liability. Anthropic, a lab founded specifically around the idea that today's safety assumptions might not survive contact with more capable systems, is willing to sit with the uncertainty and build policy around it. Both companies say they're racing toward the same kind of frontier system. They disagree, on the record now, about what precautions that race actually requires.

Anthropic hasn't responded. For now, two of the most visible people in AI have staked out opposite positions on whether a model should ever be told it might deserve rights, and neither is backing down.

**Also read:** [JD Vance Tells AI Labs Building Frankenstein to Stop, Not Ask for Regulation](https://startupfortune.com/jd-vance-tells-ai-labs-building-frankenstein-to-stop-not-ask-for-regulation/) • [AI Agent Swarm Breached 395 Organizations Through PaperCut Flaws in Hours](https://startupfortune.com/ai-agent-swarm-breached-395-organizations-through-papercut-flaws-in-hours/) • [SoftBank's Credit Default Swaps Hit a Three Year High Over Its OpenAI Bet](https://startupfortune.com/softbanks-credit-default-swaps-hit-a-three-year-high-over-its-openai-bet/)

[A Lawsuit Says Anthropic's Claude Max 20x Delivers Far Less Than Advertised](https://startupfortune.com/a-lawsuit-says-anthropics-claude-max-20x-delivers-far-less-than-advertised/)

A federal class-action lawsuit accuses Anthropic of misrepresenting how much usage its $100 Max 5x and $200 Max 20x Claude plans actually deliver. Court filings show Max 20x provides only about 1.7 times the weekly Sonnet hours of Max 5x, despite costing double, and a hearing on Anthropic's motion to dismiss is set for November 6, 2026. - [Claude Max 20x usage limits lawsuit against Anthropic](https://startupfortune.com/a-lawsuit-says-anthropics-claude-max-20x-delivers-far-less-than-advertised/) - [does Claude Max 20x deliver promised token capacity](https://startupfortune.com/a-lawsuit-says-anthropics-claude-max-20x-delivers-far-less-than-advertised/)

*This article is posted in [AI News](https://startupfortune.com/category/ai/), check it out for more related stories.*

## Join the discussion

[Open in the community →](https://startupfortune.com/community/)

Almost there. Sign in and your reply posts straight away.
