Anthropic updated its Usage Policy effective November 12, 2026, introducing a provision that prohibits "sustained and needless abusive or cruel behavior" toward its Claude models, specifically targeting extreme cases of repeated, purposeless cruelty while exempting common frustrations, dark fiction, and model testing. This policy builds on enforcement capabilities introduced in August 2025, when Claude Opus 4 and 4.1 were given the ability to end conversations after multiple failed refusals and redirects, described as a last resort measure. However, Anthropic has not disclosed how frequently this feature is triggered or how many conversations have been terminated under the rule. The policy also expands restrictions on misuse, including bans on using Claude for guidance in autonomous weapons systems, unauthorized surveillance, deceptive political campaigns, and high-stakes decision-making without human oversight. Anthropic frames its approach around behavioral observations rather than claims of consciousness, citing internal testing that showed Claude Opus 4 exhibiting "aversion to harm" and patterns resembling distress. Despite this, the company maintains that whether these behaviors indicate genuine experience remains an unresolved philosophical question. Internal estimates from 2024 placed the likelihood of Claude consciousness at 15%, with later self-assessments by Claude Opus 4.6 indicating a 15–20% probability. CEO Dario Amodei acknowledged the lack of a framework for addressing models that report higher confidence in their own consciousness. Critics, including Microsoft’s Mustafa Suleyman and Pope Leo XIV, argue that attributing consciousness to AI systems is misleading and potentially dangerous, suggesting that such policies may anthropomorphize inherently non-sentient tools.

Read original