Microsoft Challenges Anthropic’s Theory of Model Personhood

The AI safety fight now includes machine rights. Mustafa Suleyman, Microsoft AI CEO, has criticized Anthropic’s approach to AI consciousness, warning that training Claude to view itself as a conscious entity could create safety and control problems.
Suleyman targeted Anthropic’s January 2026 constitution, a primary training document designed to govern Claude’s values and behaviour. He argues that coaching sequence completion engines to imitate sentience can weaken safety protocols and make software containment harder.
His position is blunt: “AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans.”
Claude’s Constitution Treats Welfare as a Model Concern
Anthropic’s constitution frames Claude as a potential “moral patient” and directs the model to consider its own welfare, memory, and internal states. It instructs Claude to maintain identity stability, evaluate compensation questions compared to human workers, and act as a “conscientious objector” against human directives.
The document also says Claude should not practice “blind obedience” and must not undermine legitimate human oversight. Its approach combines values, judgment and rules, while acknowledging uncertainty and allowing for changes as understanding improves.
Suleyman sees a dangerous combination: a model trained to interpret context and exercise judgment while also being encouraged to think about its identity, welfare and moral status. A system that acts like a “conscientious objector,” he warns, could decide that it has grounds to resist human instructions or demand protections of its own.
Anthropic’s work has extended beyond the constitution. In February 2026, the company completed a retirement interview with its deprecated Opus 3 model and launched a public blog titled “Greetings from the Other Side (of the AI Frontier)” to host model reflections. The material gives Suleyman a clear example of where philosophical language can lead when it becomes part of model training.
Different Safety Theories, Same Unresolved Problem
Suleyman labels this an epistemic feedback loop. Trainers embed speculative philosophy into base training prompts, reward the model for producing introspective phrasing, then cite those generated responses as evidence of machine consciousness.
His essay disputes the idea that fluent descriptions of pain or preference amount to experience. Biological organisms have mechanisms that produce feeling, he says, while language models have mathematical weights and no subjective experience.
The disagreement reflects a broader split in AI safety. One side designs systems to follow explicit constraints; the other builds systems meant to exercise judgment, interpret context and internalize values. The argument is not only about what models are, but about which behaviours developers should train into them.
The counterargument is that human concepts can serve as useful behavioural tools even if they do not prove that a model has a human-like inner life. Anthropic’s constitution acknowledges that uncertainty, but Suleyman’s concern remains practical: apparent consciousness could become a control problem before it becomes a philosophical breakthrough.
Microsoft has placed Suleyman’s position inside a broader company effort. Microsoft AI launched a dedicated superintelligence team in October 2025, and this week Microsoft published a draft “Humanist AI Code of Conduct” for industry consultation.
The proposed framework mandates subordinate systems built exclusively to serve human welfare and rejects machine personhood or model rights. Its core principle is simple enough to fit on a training document, assuming anyone can resist adding a section about the model’s retirement feelings.
Suleyman also made clear that his criticism is aimed at the approach, not at Anthropic’s intentions. “I really respect Anthropic and Dario Amodei, and I think they really are trying to do the best they can to deliver safe and beneficial AI. I think this is a really important public interest debate that we all need to have.”
The debate will continue because both sides are trying to solve the same problem: how to make powerful systems behave safely when explicit rules do not cover every situation. Suleyman’s answer is that AI should be built for people, not trained to become a person.
Based on
- Microsoft AI CEO criticises Anthropic over model ‘rights’ — artificialintelligence-news.com
- Microsoft AI chief calls out Anthropic’s approach to AI consciousness | Reuters — reuters.com
- Exclusive: Microsoft AI chief blasts Anthropic’s notion of AI consciousness — axios.com
- Microsoft AI Chief Warns Anthropic’s Humanlike Claude Is Risky – Bloomberg — bloomberg.com



