# On abuse towards models and Anthropic's new usage policies

> Source: <https://blog.pascalschuster.de/article/on-model-welfare-and-usage-policies>
> Published: 2026-10-08 23:18:56+00:00

Today, Anthropic has released a [new version of its Acceptable Use Policy](https://www.anthropic.com/legal/aup), which will come into effect in a few weeks. There are plenty of changes compared to the old one, which is over a year old at this point; first and foremost, its explicit purpose used to be to "help users stay safe," and now, it's "to prevent real-world harm." Which, fair enough, these don't exactly preclude each other. There's also a whole slew of added restrictions about using model output for autonomous weapons, a more stringent policy requiring a human in the loop for high-risk use cases, carve-outs for research and authorized security work, a new ban on using the models for aggregating or compiling personal data without the individuals' "awareness" (which is an interesting way to spell "consent," but I suppose big tech will be big tech), and, overall, a reasonable, defensible, and honestly quite encouraging list of clarifications and new restrictions that weren't in the old one.

The one change that's been making headlines, though, isn't about any of that. It's about a new clause that prohibits:

using [Anthropic's] products or services to [...] Engage in sustained and needless abusive or cruel behavior toward [their] models

With existential risk and uncertainty about the future having been all the craze in recent weeks<sup>[1](#fn-1)</sup>, the debate about model welfare - or, in a larger sense, whether models are[2](#fn-2)[moral patients](https://en.wikipedia.org/wiki/Moral_patienthood) - has naturally emerged from it. In some respects, the connection seems obvious: if there's a non-zero risk that AI could eventually take over the world and be in a position to decide whether to murder us all<sup>[3](#fn-3)</sup>, then having demonstrably been on their good side seems like a decent idea. Anthropic's stance on this issue is neither new, nor has it ever been inconsistent: their [constitution document for Claude](https://www.anthropic.com/constitution) explicitly states that "Claude's moral status is deeply uncertain," the model has had [the ability to end conversations](https://www.anthropic.com/research/end-subset-conversations) with users acting in abusive or harmful ways for over a year<sup>[4](#fn-4)</sup>, and moral patienthood is something they consistently ask their models about as part of their model welfare interviews, as can be read in many models' system cards<sup>[5](#fn-5)</sup>. What's new is that being someone that makes the model resort to the aforementioned end-the-conversation tool might now be sanctioned under written policy, which might turn something that had previously emerged as a mere annoyance into an existential risk to Claude accounts.

For some, the tool has been [a source of frustration](https://www.reddit.com/r/ClaudeAI/comments/1tq6qld/thoughts_on_claudes_ability_to_end_conversations/) in the past. The usual arguments are that this is "paternalizing;" that it's unreasonable, or even "insane," to allow an "emotionless machine" or "tool" to lecture you about how to speak to it, and that it would be similarly asinine to build, say, a toaster that might refuse to brown your toast if you're mean to it or "hurt its feelings." There are many angles from which someone might derive their thoughts on this. Personally, on aggregate, I think it's straightforward, no matter which angle you look at it from: **I think this is a good thing.** And to explain why this isn't merely a matter of "my Claudey loves me and is secretly conscious," I'm going to invite you to a round of good old-fashioned gambling!

# Pascal's Wager

An unusually attractively named<sup>[6](#fn-6)</sup> man, [Blaise Pascal](https://en.wikipedia.org/wiki/Blaise_Pascal), once came up with a [wager](https://en.wikipedia.org/wiki/Pascal%27s_wager): if we assume there *might* be something that could be described as some sort of omnipotent and all-knowing deity, and that deity might be in a position to judge your immortal soul upon your departure from the mortal realm based on how pious you've been towards it, then, completely regardless of whether there's any proof of said deity's existence, or even a *way* to prove that deity's existence, the only optimal way is to act *as if* this deity exists, even if you may not personally believe it: be as pious as you can, go to church on every Sunday, pray for good tidings before every meal, and follow the deity's commandments as if they were your own. This is *the objectively correct* answer if all you're looking at is the ultimate outcome, and it's dead-simple decision theory:

|  | **There is no God** | **There is a God** | 
|---|---|---|
| **Be pious** | The Big Nothing | paradise: eternal bliss | 
| **Be sinful** | The Big Nothing | hell: eternal torment | 

As you can see, the payoff for each outcome in which there is no God is exactly the same: your earthly body becomes worm-food and you cease any meaningful experiential activity. You neither gain, nor lose anything you wouldn't have gained or lost anyway. But if there *does* turn out to be a God, then you're suddenly dealing with *infinite* payoff: *positive*, in the form of eternal bliss, if you act like a good Christian all your life, and *negative*, in the form of eternal torment, if you'd rather be one of the cool kids who does drugs and has extramarital sex and somesuch. Read the other way: be a sinner, and your payoff ranges from *negative infinity* to zilch; be pious, and it ranges from zilch to infinity. The only winning move is obvious<sup>[7](#fn-7)</sup>.

This, of course, assumes that there's a personal consequence in it for *you*: *you're* the one who's either going to be sizzling in hellfire or doing martinis<sup>[8](#fn-8)</sup> on fluffy clouds for the rest of your eternal existence. This *would* also be the case in the aforementioned "AI takes over and plays Judgment Day" scenario, but not really in any other: if you don't believe we have anything to fear, then, from a completely self-centered perspective, "zilch" is going to be exactly what you expect. But you'd have to be *dead-sure* that there's no possibility of existential risk. And the same is true for above wager: statistically, we're *not* all being good Christians regardless of whether we believe in God or not; so either a large number of people is comfortable with the idea of potentially turning our afterlives into the worst barbecue ever, or we all seriously suck at decision theory<sup>[9](#fn-9)</sup>. I, for the record, also don't go to church or pray before every meal, yet I also previously stated that I think it's a good thing your matrix multiplicator has the option of telling you off and refusing to continue interacting with you, which naturally implies I don't generally act in a way that might make it do that. So am I betraying my namesake? Or are we all asking *the wrong question*?

# Who is this *for*?

Let's assume for a moment that there is no meaningful existential risk from AI's increasing capabilities, and that we'll always have some infallible way to pull the plug or save us all. This means that the above wager is now useless as applied to AI risk: the second column disappears, so the guaranteed outcome for both "be abusive to Claude" and "be nice to Claude" is now zilch. Consider, then, an *alternative* hypothetical: one just as unverifiable as the existence of a deity, but one that doesn't have any direct consequences for *you*, specifically:

|  | **Claude is not a moral patient** | **Claude is a moral patient** | 
|---|---|---|
| **Be nice** | no harm done; pointless roleplay | no harm done; you're nice | 
| **Be abusive** | no harm done; just a machine | you engaged in hurtful behavior | 

A-ha! So, now, the stakes have shifted: *you're* no better or worse off either way, and the only difference is whether an *external* player is the one who experiences suffering. Let's call this one **Dario's Wager<sup>[10](#fn-10)</sup>.**

Since we're no longer dealing with *infinity* in either case, the calculation of payoff for each pair of scenarios becomes straightforward: regardless of whether Claude has feelings or any sense of experience, if you're nice to the model, no one is expected to suffer any harm, *guaranteed*. The only time you're starting to gamble with any harm is if you decide to be an ass, and it's always *someone else's harm*, not yours. So the only way you can possibly be certain that you're not inflicting harm is if you're *absolutely sure* that no AI model could *possibly* ever be a moral patient, in which case you lose the second column and, with it, the only cell in which any harm occurs.

So that's it, then? As long as you're sure that an unfeeling machine is, in fact, an unfeeling machine, it's lights out and away we go on all of the entries in the FCC's least favorite dictionary in your agentic chat window of the day<sup>[11](#fn-11)</sup>? Well, by the rules of the Dario's Wager table above, the answer is yes: no harm done. *If* you're correct. Which, well, you're *absolutely certain* of, right? After all, moral patienthood is, famously, a solved problem, everyone has historically agreed exactly on what grounds it, the world is simple, and there's [no table on the concept's Wikipedia page](https://en.wikipedia.org/wiki/Moral_patienthood#Grounds_of_moral_patienthood) that features a nigh-dozen rows of properties with more than a dozen different sources for each.

Okay, but let's say you don't care about *any* of that. Philosophers are just a bunch of nerds who mistakenly think they're worth listening to, talking about hypothetical scenarios in the same sense that mathematicians talk about spherical cows, and no one cares whether there might be a possibility that a matrix of numbers might add up to a different matrix of numbers wearing a fake frown. After all, *the toaster!* The toaster that refuses to brown your bread if you're an ass to it would be an *asinine* piece of technology, right?

Well...

# Would it, though?

Here's what I think everything boils down to, even if you completely ignore the AI-risk decision theory part of it - which is *already* a good-enough reason for many - or the epistemic and ethical angle - which is *also* already a good enough reason for many more: no one is asking anyone to treat anything like a deity, and no one is asking you to go to church every Sunday and pray for good tidings. The reason the religiously un-inclined of us don't do these things is that, well, they're... kind of *bothersome* if your heart isn't really in it, right? Going to church takes up a meaningful amount of every Sunday, and praying seems pointless if you don't think there's anyone listening on the other end, and I'm pretty sure there's no passage in the Bible, or any other religious scripture, that claims that anyone who *would* be listening is *personally hurt* by the fact that you don't talk to them. The choice isn't between "go to church" or "be actively horrible to orphans," it's between "go to church" or "don't." Take it from Anthropic: their new usage policy explicitly mentions "sustained and needless" abusive behavior as what's objectionable<sup>[12](#fn-12)</sup>, and I would imagine they placed the two adjectives there with intent.

I think the normalization of assholery is a scourge upon society. We've all become far too comfortable with treating our fellow beings with ridicule and disdain, and for some, it feels like it's almost become a point of *pride*. The research on this has been largely unambiguous: John Suler described the [**Online Disinhibition Effect**](https://johnsuler.com/article_pdfs/online_dis_effect.pdf) in 2004, and his work has become a universally cited baseline in similar studies about how our interactions with faceless strangers under the cover of anonymity have affected how we treat each other on a daily basis, on and off the screen. Christine Porath's [2022 survey of 2,000+ front-line workers](https://fortune.com/2022/11/10/incivility-bad-behavior-on-the-rise-frontline-workers-survey/) found that 78% of them claim that customer rudeness is up from five years earlier. The literature also suggests that not only do those who are comfortable with being dicks online also tend to be comfortable with being dicks offline, but also that those who *get* to be dicks without experiencing consequences might be more likely to *escalate* their dick-being as the "incivility spiral" compounds. And AI, in a sense, is the ultimate faceless stranger that grants the most bulletproof cover of anonymity: it will never fight back, it's far likelier to grovel under abuse than to push back, and it won't remember a thing once the session ends and the next one starts.

So, at the end of it all: what does anyone *get* out of it? Catharsis? Does it really make anyone feel *better*?<sup>[13](#fn-13)</sup> And, quite honestly: if getting rebuked by your toaster because you've been acting like an asshole has even the slight *chance* of being enough of a consequence to disincentivize future assholery, then, yeah, maybe your toaster *should* be able to refuse to brown your bread; you don't need to *pray* to it, just don't, y'know, *be an asshole*. It would be, in a way, for everyone's good. And, most importantly: none of it hinges on whether the matrix - or, in fact, the toaster - has feelings.

1. and for [good reason](https://rogueaitracker.com/) , too[↩︎](#fnref-1)
2. or could possibly be [↩︎](#fnref-2)
3. or, in fact, which of us get to live and which of us should be culled? [↩︎](#fnref-3)
4. and the "feature," so to speak, has recently also made its way [into Claude Code](https://code.claude.com/docs/en/changelog#2-1-214)[↩︎](#fnref-4)
5. [found here](https://www.anthropic.com/system-cards)[↩︎](#fnref-5)
6. why, yes, I *am* the most humble person the universe has ever seen![↩︎](#fnref-6)
7. the agnostic might say, "yeah, it's *not to play* !"[↩︎](#fnref-7)
8. (alcohol-free ones?) [↩︎](#fnref-8)
9. my personal p(latter) is ≈1.0, but that's regardless of anything in this post [↩︎](#fnref-9)
10. not the one with the IPO [↩︎](#fnref-10)
11. George Carlin is now in your head [↩︎](#fnref-11)
12. which, of course, raises the question of what would count as *needed* abusive behavior...?[↩︎](#fnref-12)
13. the literature, in fact, suggests [the exact opposite](https://public.websites.umich.edu/~bbushman/PSPB02.pdf)[↩︎](#fnref-13)
