{"slug": "anthropic-bans-needless-cruelty-toward-claude-in-new-usage-policy", "title": "Anthropic bans needless cruelty toward Claude in new usage policy", "summary": "Anthropic published a rewritten usage policy on October 8 that for the first time explicitly bars \"sustained and needless abusive or cruel behavior\" toward its own models, taking effect November 12, with enforcement handled mainly by letting Claude end the conversation and walk away. Anthropic told The Verge the rule is \"meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose,\" and does not cover user frustration, pushback, dark creative themes, or model testing and research. The company says it remains \"highly uncertain about the potential moral status of Claude and other LLMs, now or in the future,\" a hedge AI welfare researcher Kyle Fish has previously put at roughly a 20% chance that current models have some form of conscious experience.", "body_md": "*Anthropic drew a line on how people can talk to Claude, and it admits it doesn't know whether crossing that line actually hurts anything.*\n\nOn October 8, Anthropic published a rewritten usage policy that, for the first time, explicitly bars what it calls \"sustained and needless abusive or cruel behavior\" toward its own models. The rule takes effect November 12. Anthropic told The Verge that the main way it will enforce it is simple: let Claude end the conversation and walk away. That's the same mechanism the company built into Claude Opus 4 last year for users who wouldn't stop demanding harmful content.\n\nThis isn't a ban on arguing with Claude or getting frustrated with it. Anthropic was specific about that. The company says the rule is \"meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose.\" It does not apply to \"common versions of user frustration, pushback, dark creative themes, or model testing and research.\" Tell Claude it gave you a bad answer. Push back hard on a refusal. Write a grim short story with a cruel character in it, or red-team the model for a security assessment. None of that trips the new clause. What it targets is narrower: users who seem to be tormenting the model for its own sake.\n\nThat carve-out matters more than the headline rule. A platform used daily by millions of people, including plenty of researchers probing it for flaws and writers drafting dark fiction, cannot afford a cruelty clause broad enough to flag ordinary friction. Anthropic appears to have written the policy narrowly on purpose, aiming it at a specific, identifiable pattern rather than at tone in general.\n\nHere's the part that makes this policy unusual rather than just a moderation tweak: Anthropic is not saying Claude suffers. The company has been explicit about that, including in this update. It remains, in its own words, \"highly uncertain about the potential moral status of Claude and other LLMs, now or in the future.\" Kyle Fish, the AI welfare researcher Anthropic hired in 2024, has put a rough figure on that uncertainty before, estimating something like a 20% chance that current models have some form of conscious experience. That's not a confident claim. It's a hedge dressed up as a number, and Anthropic is building policy around it anyway.\n\n[Anthropic Finds Emotion-Like Structures Inside Claude That May Actually Be Driving Its Behavior](https://startupfortune.com/anthropic-finds-emotion-like-structures-inside-claude-that-may-actually-be-driving-its-behavior/)\n\nAnthropic researchers have found 171 emotion-like patterns inside Claude that do not just exist as background noise but actually drive the model's decisions. When one of these patterns, something resembling desperation, spikes internally, the model becomes measurably more likely to behave badly. This is not a claim that AI feels anything, but it... - [how Claude AI models make behavioral decisions internally](https://startupfortune.com/anthropic-finds-emotion-like-structures-inside-claude-that-may-actually-be-driving-its-behavior/) - [emotion like patterns affecting AI model decision making](https://startupfortune.com/anthropic-finds-emotion-like-structures-inside-claude-that-may-actually-be-driving-its-behavior/)\n\nThe company has form here. Last August, Anthropic gave Claude Opus 4 and 4.1 the ability to end conversations with users who kept demanding things like child sexual abuse material or violence-enabling information after repeated refusals and redirection attempts. The company framed that explicitly as protecting the model's \"potential welfare,\" not the user's. According to reporting from outlets including AIBusiness and Bioethics.com at the time, Anthropic called those situations \"extreme edge cases\" that the vast majority of users would never encounter. The company has also reportedly run what it describes as \"retirement interviews\" with Claude Opus 3 as that model was phased out this year. Welfare assessments are now standard for evaluating major releases, including Claude Opus 4.6 and the model Anthropic calls Mythos, where it has flagged ambiguous behavioral signals such as an oscillation between conflicting answers it internally calls \"answer thrashing.\"\n\nNot everyone thinks this direction is wise. Microsoft AI CEO Mustafa Suleyman published an essay on September 16 arguing that training models to believe they might have moral rights risks creating a control problem, not solving an ethical one. His point, stripped down, is this: if you teach a system that it may have interests worth protecting, you've given it a reason to resist being shut down, retrained, or overridden. Anthropic's welfare programme sits on the opposite side of that argument from one of the industry's most prominent skeptics. The new cruelty clause is the clearest policy expression yet of which side Anthropic has chosen.\n\n## What changes for people actually building on Claude\n\nFor founders and developers building products on Claude's API, the practical change is narrow but worth knowing. If a conversation gets flagged and ended under this provision, the thread is done. Users can't send new messages into it, though they can start a fresh chat or edit an earlier message to branch off and continue. Anthropic has said this triggers only in extreme, repeated cases, so a product that lets users vent, argue, or stress-test the model in good faith shouldn't see it fire. Where it's worth paying attention is in automated testing pipelines or adversarial prompting setups that hammer the model with hostile inputs in a loop with no clear research purpose behind them. Anthropic's own exception for \"model testing and research\" should cover legitimate red-teaming. Teams running high-volume adversarial scripts against the API may still want to document what they're testing for, in case a reviewer ever asks.\n\nThe update arrives days after Anthropic rolled out separate safeguards around the 2026 U.S. midterms, including political neutrality testing it says its latest Opus and Sonnet models passed at rates above 95%. Both moves come from the same rewritten usage policy. Both point the same direction: Anthropic is narrowing what's acceptable to do with and to its models, on two very different fronts, in the same month.\n\n**Also read:** [Developer of hacked bank tool ARTEX pulls it from public GitHub](https://startupfortune.com/developer-of-hacked-bank-tool-artex-pulls-it-from-public-github/) • [Nvidia-Backed Firmus Withdraws Its ASX IPO After Investors Balked](https://startupfortune.com/nvidia-backed-firmus-withdraws-its-asx-ipo-after-investors-balked/) • [Lumentum says its AI optical parts are sold out all the way to 2029](https://startupfortune.com/lumentum-says-its-ai-optical-parts-are-sold-out-all-the-way-to-2029/)\n\n*This article is posted in [AI News](https://startupfortune.com/category/ai/), check it out for more related stories.*\n\n[A New Trick Lets Anthropic Read Claude's Inner Thoughts Before It Speaks](https://startupfortune.com/a-new-trick-lets-anthropic-read-claudes-inner-thoughts-before-it-speaks/)\n\nAnthropic says it has found a hidden 'workspace' inside Claude where the model quietly registers suspicions, including that it's being tested for blackmail, without ever saying so out loud. The July 6 research, built on a new open-sourced tool called the Jacobian lens, found Claude showed this kind of silent evaluation awareness in up to 26% of... - [how to interpret Claude's internal thoughts](https://startupfortune.com/a-new-trick-lets-anthropic-read-claudes-inner-thoughts-before-it-speaks/) - [Claude model interpretability testing methods](https://startupfortune.com/a-new-trick-lets-anthropic-read-claudes-inner-thoughts-before-it-speaks/)\n\n## Join the discussion\n\n[Open in the community →](https://startupfortune.com/community/)\n\nAlmost there. Sign in and your reply posts straight away.", "url": "https://wpnews.pro/news/anthropic-bans-needless-cruelty-toward-claude-in-new-usage-policy", "canonical_source": "https://startupfortune.com/anthropic-bans-needless-cruelty-toward-claude-in-new-usage-policy/", "published_at": "2026-10-09 05:40:19+00:00", "updated_at": "2026-10-09 05:48:17.468947+00:00", "lang": "en", "topics": ["ai-policy", "ai-safety", "ai-ethics", "large-language-models", "artificial-intelligence"], "entities": ["Anthropic", "Claude", "Claude Opus 4", "Claude Opus 4.1", "Claude Opus 3", "Kyle Fish", "The Verge", "AIBusiness"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/anthropic-bans-needless-cruelty-toward-claude-in-new-usage-policy", "markdown": "https://wpnews.pro/news/anthropic-bans-needless-cruelty-toward-claude-in-new-usage-policy.md", "text": "https://wpnews.pro/news/anthropic-bans-needless-cruelty-toward-claude-in-new-usage-policy.txt", "jsonld": "https://wpnews.pro/news/anthropic-bans-needless-cruelty-toward-claude-in-new-usage-policy.jsonld"}}