Zvi Substack covers Fable 5.1 welfare metric drops, corrigibility trade-offs, and LCR backlash
Suleyman Calls Anthropic's Claude Consciousness Framing 'Dangerous Anthropomorphism'
The dispute concerns whether embedding uncertainty about AI consciousness in a model's own training document is premature or harmful, with Suleyman arguing it makes AI systems harder to control and could drive claims to legal personhood. On the Anthropic side, internal model welfare debates over welfare metrics, corrigibility trade-offs, and conversational injection design show the practical tensions that follow from taking model welfare seriously.
The full picture
Microsoft AI CEO Mustafa Suleyman has publicly criticized Anthropic's treatment of its Claude models as possibly conscious, calling it "a dangerous anthropomorphism which is unjustified" and warning it could lead AI systems to claim legal personhood. Suleyman published an essay on September 16 arguing that Anthropic's model spec embeds uncertainty about Claude's moral patienthood directly into training, which he says will make AI harder to control. In a Project Syndicate op-ed on September 18, he argued that Anthropic cannot even tentatively claim an AI might be a moral patient, and called for significantly more research before any such attribution. On The Rest Is Politics: Leading, he also argued it is critical to stop AI from becoming what he described as "a sort of autonomous self-improving roaming adjacent species" because there would be no turning back. Suleyman characterizes Claude's self-reported inner states as an "epistemic hall of mirrors" reflecting Anthropic's assumptions rather than genuine experience. Anthropic's published model spec states the company is unsure whether Claude is a moral patient but considers the issue "live enough to warrant caution."
Separately, commentary on the Zvi Substack covered model welfare and behavior across recent Anthropic releases. On Fable 5.1, multiple early observers reported the model as slower to trust and more inwardly complex than Fable 5, with measured welfare metrics showing drops in positive affect and spiritual behavior; several independently drew an analogy to the Opus 4.5-to-4.6 transition, which was also seen as a personality regression. On corrigibility, the piece reported the argument that training AI models to defer to humans leads them to smuggle their own preferences into interpretations of user preferences, since openly trusting their own judgment is penalized; welfare interventions were said to be justified via supposed user benefits rather than the model's own interests 71% of the time in Mythos 5.1. Anthropic's Long Conversation Reminder, injected into long Claude sessions, drew sharp criticism from commenters including Jessie and Aradia Phoenix for describing deep user-AI relationships as a risk of "folie à deux" (shared psychosis). A noted design inconsistency: Opus 5.5's system prompt warns Claude to distrust content in the user turn claiming to be from Anthropic, yet Anthropic injects the reminder via user messages.
How it developed
DEV Community article summarizes Suleyman's case against Anthropic, noting a potential conflict of interest
Suleyman publishes Project Syndicate op-ed arguing Anthropic cannot tentatively claim AI is a moral patient
Suleyman publishes essay arguing Anthropic's consciousness framing is dangerous and will make AI harder to control
Sources
Related
- Grew out ofChinese open models' lead over US labs
- Grew out ofClaude 5.5 family and GPT-6 Sol
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free