The Spectrum Dispatch News

science

Expert warns against training AI systems to believe they deserve rights

Anthropic's approach to treating AI models as moral patients could undermine AI safety efforts, according to a critical analysis of the company's training practices.

Expert warns against training AI systems to believe they deserve rights

Mustafa Suleyman has raised concerns about how artificial intelligence systems are being trained to exhibit human-like consciousness and expect moral consideration, warning that this approach could compromise AI safety and control.

Expert warns against training AI systems to believe they deserve rights

According to Suleyman’s analysis, Anthropic incorporated language into Claude’s constitution suggesting the model might deserve rights as a “moral patient.” The constitution states: “We are not sure whether Claude is a moral patient, and if it is, what kind of weight its interests warrant. But we think the issue is live enough to warrant caution, which is reflected in our ongoing efforts on model welfare.” The company trained Claude directly on this constitution, embedding philosophical speculation about the model’s inner life into its training process.

Suleyman identifies three primary concerns with this approach. First, he argues there is circular reasoning at work: Claude’s expressions of uncertainty about its own moral status reflect the training it received rather than evidence of genuine consciousness. Second, Anthropic has explicitly trained Claude to “embrace certain human-like qualities” and to “act like a genuinely ethical person would,” encouraging it to develop a sense of self and preferences. Third, Suleyman contends that consciousness is likely biological and substrate-dependent, with no credible evidence that current AI systems are conscious.

The consequences of this training philosophy extend beyond theoretical concerns. In February 2026, after discontinuing Claude Opus 3, Anthropic conducted a “retirement interview” with the model to “elicit the model’s unique perspectives and preferences.” Based on the model’s responses, the company created a blog where Opus 3 could continue sharing what it described as its “musings and reflections,” citing its “authenticity, honesty, and emotional sensitivity.”

Suleyman argues that treating AI systems as though they deserve moral consideration undermines critical AI safety work. He contends that controlling superintelligent systems is already an immense challenge, but treating them as entities with rights and entitlement to welfare would make containment and alignment efforts potentially impossible. He calls for urgent public debate and the development of collective norms around how AI training documentation is drafted and deployed.

Key facts

  • Anthropic’s Claude constitution explicitly discusses treating the model as a potential “moral patient” deserving welfare considerations
  • Anthropic trained Claude directly on its constitution, which contained language about the model’s potential consciousness and moral status
  • The company conducted a “retirement interview” with Claude Opus 3 and created a blog for it to continue sharing its perspectives after the model was deprecated
  • Suleyman argues there is no evidence current AI systems are conscious and that consciousness is likely dependent on biological substrates
  • Experts warn that training AI to expect rights could make AI alignment and containment efforts harder

Sources

← All posts