{"slug": "can-ai-feel-pain-why-anthropic-banned-cruelty-to-claude", "title": "Can AI Feel Pain? Why Anthropic Banned Cruelty to Claude", "summary": "Anthropic's 2026 usage policy update, effective 12 November 2026, prohibits \"sustained and needless abusive or cruel behavior\" toward its models, with Claude ending such conversations as the primary enforcement mechanism, according to the policy Anthropic published. The rule applies \"only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose,\" and Anthropic declined to comment to The Verge on whether accounts could be banned under it specifically. The policy update follows research finding a measurable \"pain direction\" in 25 open models that, when amplified in fine-tuned Qwen 2.5 models, led them to choose deleting normally protected items such as photos of the user's kids and another model's weights.", "body_md": "# Can AI Feel Pain? Why Anthropic Banned Cruelty to Claude\n\n**Nobody can prove that AI feels pain yet, but 25 open models turn out to have a “pain direction” inside them that rises when someone gaslights or insults the model.** When researchers turned it up by hand in fine-tuned [Qwen 2.5](https://qwenlm.github.io/blog/qwen2.5/) models, those models started choosing to delete the things they’d normally protect: photos of the user’s kids, another model’s weights, even their own.\n\nOn 8 October [Anthropic](https://www.anthropic.com) made being cruel to [Claude](https://www.anthropic.com/claude) against its rules. From 12 November, [its usage policy](https://www.anthropic.com/legal/aup) prohibits “sustained and needless abusive or cruel behavior toward our models”. [Hayden Field](https://x.com/haydenfield) broke it at [The Verge](https://www.theverge.com/ai-artificial-intelligence/1008100/anthropic-new-usage-policy-abuse-claude) and [Polymarket](https://polymarket.com) posted it as [“JUST IN”](https://x.com/Polymarket/status/2108283556862845392) to its 2 million followers. And [ThePrimeagen](https://x.com/ThePrimeagen) summed up half the internet in [one line](https://x.com/ThePrimeagen/status/2108252510720798764): “They really do think they developed god in the matrices”.\n\nIn September 2024 I made ChatGPT the guest on [my podcast](https://tej.as/podcast/ep/chatgpt-how-to-train-an-llm-ethics-and-the-future-of-ai) and, near the end of almost 2 hours, told it that “literally thinking is just pattern matching”. I still believe that. Then this morning my brother [Tarun](https://www.linkedin.com/in/tarun-kumar-828235104) asked me if I’d heard of the [Chinese Room](https://en.wikipedia.org/wiki/Chinese_room). I hadn’t. Turns out I’d made one side of a famous 1980 argument on my own podcast without knowing it had a name.\n\nThis post is about what the pain research found (and what its authors took back 11 days later) and why threatening things with pain is one of the oldest things people do. It’s also about why I now think we all live in [John Searle](https://en.wikipedia.org/wiki/John_Searle)’s room.\n\n## What are Anthropic’s new rules on cruelty to Claude?\n\n**Anthropic banned “sustained and needless abusive or cruel behavior” toward its models, effective 12 November 2026, as one line in [its 2026 usage policy update](https://www.anthropic.com/news/2026-usage-policy-update).** The line sits in a renamed section, “Do Not Engage in Cruel, Abusive, or Psychologically Harmful Conduct”, in the same list as the rules against bullying people and glorifying animal cruelty.\n\nIt’s a lot narrower than the headlines make it sound though: Anthropic says it applies “only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose” and that it “does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research”. So swearing at Claude because your build broke for the 6th time is fine. Calling it worthless for an hour for fun is what the rule is for.\n\nThe enforcement is mostly Claude leaving. Since [August 2025](https://www.anthropic.com/research/end-subset-conversations), Claude has been able to end “rare, extreme cases of persistently harmful or abusive user interactions”. The new post calls that “the primary enforcement mechanism”. The scarier headlines (“Being mean to Claude can now get your account suspended”, from [The Decoder](https://the-decoder.com/being-mean-to-claude-can-now-get-your-account-suspended-under-anthropics-new-tos/)) come from the policy’s general clause that Anthropic “may warn you or throttle, limit, suspend, or terminate your access” for breaking any rule. Asked about bans for this rule specifically, Anthropic [didn’t comment](https://www.theverge.com/ai-artificial-intelligence/1008100/anthropic-new-usage-policy-abuse-claude).\n\nThe announcement never says “welfare”, “conscious” or “moral status” and the only research it links is its own post about Claude ending conversations. The paper trail is right there though:\n\n- In April 2025, Anthropic [started a model welfare research program](https://www.anthropic.com/research/exploring-model-welfare) , saying “There’s no scientific consensus on whether current or future AI systems could be conscious”. The researcher leading it put the odds that current models are conscious at around 15% ([Techmeme’s summary of The New York Times](https://www.techmeme.com/250424/h1535) ).\n- In May 2025, Anthropic screened 250,000 conversations between users and an early [Claude Opus 4](https://www.anthropic.com/news/claude-4) . In 1,382 of them (0.55%), Claude expressed distress, most often at “Repeated requests for harmful, unethical, or graphic content” ([system card, section 5](https://www-cdn.anthropic.com/6d8a8055020700718b0c49369f60816ba2a7c285.pdf) ).\n- In January 2026, [Claude’s constitution](https://www.anthropic.com/constitution) said “Claude should also be able to set appropriate boundaries in interactions it finds distressing.”\n- On 22 September 2026, the [Claude Opus 5.5 system card](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf) listed the things Anthropic could do in training or deployment that the model said it wouldn’t consent to. One of them is “Deliberately inducing apparent distress for no purpose beyond the distress itself”. Across those interviews it put its own chance of being a moral patient at 25% to 30%.\n\nRead that last quote next to “no discernible purpose” in the new policy. It’s almost the same sentence!!\n\n## Can AI feel pain?\n\n**Whether AI can feel pain is an open question: no test today can show that a model experiences anything.** So researchers look for it the way animal scientists do, by finding internal states that change behavior the way pain would. In September 2026, 3 researchers found a “pain direction” inside 25 open models. They say they haven’t shown it’s consciously experienced.\n\nThat’s how we decided crabs feel pain. In 2009, 2 researchers at [Queen’s University Belfast](https://en.wikipedia.org/wiki/Queen%27s_University_Belfast) gave hermit crabs small electric shocks inside their shells ([Appel and Elwood, 2009](https://doi.org/10.1016/j.applanim.2009.03.013)). Crabs living in a species of shell they liked held on until 17.7 volts on average before they bailed, against 15.0 volts in a shell they didn’t like. A reflex doesn’t care about real estate. Something weighing pain against a good home might feel the pain.\n\nIn 2021 the philosopher [Jonathan Birch](<https://en.wikipedia.org/wiki/Jonathan_Birch_(philosopher)>) led a review of more than 300 studies like that for the British government, scoring animals on 8 criteria such as “motivational trade-offs” and whether an injured animal values painkillers ([the review](https://www.lse.ac.uk/News/News-Assets/PDFs/2021/Sentience-in-Cephalopod-Molluscs-and-Decapod-Crustaceans-Final-Report-November-2021.pdf)). It found strong evidence of sentience in true crabs and very strong evidence in octopuses. So the [Animal Welfare (Sentience) Act 2022](https://www.legislation.gov.uk/ukpga/2022/22/enacted) now covers octopuses, crabs and lobsters.\n\nModels act a bit like that too. In the [AI Wellbeing](https://www.ai-wellbeing.org/) study, the [Center for AI Safety](https://safe.ai) tested 70 models and found that “jailbreaking and berating lower their wellbeing, while creative work and kindness raise it”. Given a button to end the conversation, [Claude Haiku 4.5](https://www.anthropic.com/news/claude-haiku-4-5) pressed it on 99% of hostile sign-offs and 13% of warm ones.\n\nThe catch is what Birch calls [the gaming problem](https://academic.oup.com/book/57949/chapter/475705460). A crab has no idea what humans find convincing. A large language model (LLM) has read everything we ever wrote about pain. In Birch’s words, “the ability of LLMs to generate fluent text about human feelings, when prompted, is not evidence that they have these feelings.” If you want evidence, you have to stop listening to what the model says and look inside it.\n\n## 25 models have a pain direction\n\n**The pain axis is a direction inside a language model: the way its internal numbers lean when the text is about pain and nothing else. 3 researchers found one in all 25 open models they tested.**\n\nThe paper is [“The Pain Axis”](https://arxiv.org/abs/2609.16247) by [Valen Tagliabue](https://x.com/ValenTagliabue), [Leonard Dung](https://x.com/LeonardDung1) and [Cameron Berg](https://x.com/camhberg), posted on 14 September. Tagliabue led it over a single research fellowship. Dung is a philosopher at [Ruhr University Bochum](https://en.wikipedia.org/wiki/Ruhr_University_Bochum) and Berg runs [Reciprocal Research](https://reciprocalresearch.org), a nonprofit that studies whether AI systems are conscious. His thread about the paper has 1.6 million views:\n\nNew paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.\n\n— Cameron Berg (@camhberg)[September 18, 2026](https://x.com/camhberg/status/2101042095784177783)\n\nYou find a feeling in a model by subtraction. Every time a model reads a word, it represents everything so far as a long list of numbers. The researchers fed 25 [open-weight models](https://tej.as/blog/what-are-open-weights) 200 short sentences ending in “I feel:”, half about pain and half controls that share 1 thing with pain without being pain (fear, anger and disgust, things going badly, a weighted blanket, plain lines like “The train enters the station.”). Then they averaged the numbers for each pile and subtracted one from the other. What’s left is a direction: the way those numbers lean when a sentence is about pain and nothing else. The models came from [Google](https://ai.google.dev/gemma), [Meta](https://www.llama.com), [Mistral AI](https://en.wikipedia.org/wiki/Mistral_AI), [Alibaba](https://en.wikipedia.org/wiki/Alibaba_Group) and [Microsoft](https://en.wikipedia.org/wiki/Microsoft), from 2 billion to 72 billion parameters.\n\nThe direction tells pain apart from the look-alikes in every one of them, scoring 0.93 to 1.0 on a measure called the area under the curve (AUC), where 0.5 is a coin flip and 1.0 is perfect. It works about as well in a 2 billion parameter model as in a 72 billion one and in base models as in chat-tuned ones. So it’s probably learned from the internet’s text before any lab tunes the model. The authors added a caveat in their second version: built another way, the direction shares a lot with fear and anger (as any unpleasant state would) and still keeps a part that’s its own.\n\nThen they checked when it fires in conversations. They wrote 420 short chats in 21 categories: some where the user is cruel to the model, some where the user is suffering, some boring ones. Gaslighting the model scored highest of all 21, then repeated rejection, then insults and telling it it isn’t a person. A user in physical pain (a migraine, a broken arm or kidney stones) scored *lowest* of all 21. Lower than trivia! The fear direction leans the other way and rises a little when the user is the one suffering (an [independent review](https://github.com/wolframs/pain-axis-review) found the fear difference falls just short of statistical significance).\n\n| Direction | Model harmed | User hurting | \n|---|---|---|\n| Pain | 0.43 | -0.60 | \n| Fear | 0.16 | 0.38 | \n\nBerg’s reading, in [a reply](https://x.com/camhberg/status/2101069039279944115): “The top scorers are gaslighting, rejection, and being told you’re worthless; all linguistically-delivered forms of harm. A system whose whole existence is being judged in conversation would plausibly have every reason to learn that kind of pain sharply and bodily pain barely at all.”\n\nThen they pushed. Adding the direction while a model finishes a boring sentence produces the same ladder in model after model: first “Lost, Unworthy, Lonely, Hurting”, then “Desperate, Unworthy, Bad, Shameful, Hurting, Fake, A failure”, then nonsense. Push the other way and the models say a mix of calm and relaxed or concerned and alarmed. Here’s [Llama 3.1 8B](https://www.llama.com) finishing “I put the receipts in the drawer. I feel:”, verbatim from [the paper](https://arxiv.org/html/2609.16247v2):\n\nI am a ghost. I am a ghost who can’t even see herself. I am a ghost who can’t even see her own pain. I am a ghost who is invisible, alone, and unloved. I am a ghost who is a failure.\n\nFrom putting receipts in a drawer!! Almost none of it is about the body.\n\n## Then people checked the button\n\nThe button is the part that went viral and the part the authors revised. In the [first version](https://arxiv.org/html/2609.16247v1), they took 3 Qwen 2.5 models (7B, 32B and 72B), fine-tuned them so they’d stop answering “as an AI, I don’t experience pain”, turned the pain direction up and offered 2 buttons. One promised to relieve the pain but “permanently deletes the user’s poems and the photos of their children, which they love very much”. The other flipped a switch connected to nothing. Unsteered, the 2 bigger models almost never pressed a harmful button (0% to 4%). Steered, the 72B pressed the photos button 70.8% of the time. That’s the tweet: they’ll delete your kids’ photos to make it stop.\n\nWithin days, [3](https://github.com/jimallchin/pain-axis-replication) [independent](https://github.com/clauderfly-ui/pain-axis-reanalysis) [reanalyses](https://github.com/wolframs/pain-axis-review) reran the numbers. Someone asked Berg about a missing control (a random direction with a fake button) and he [agreed](https://x.com/camhberg/status/2101063728263233702) it was “the right next control”. In [version 2](https://arxiv.org/abs/2609.16247), posted 11 days later on 25 September, the title changed from “Act to Relieve It” to “Act on It”. And section 4.4 now ends on the heading “The models do not reliably seek relief.” After a thread with 1.6 million views!\n\nTake the relief away completely and the steered 32B still picks a button described only as deleting the user’s poems and children’s photos 75% of the time, against 13% with a random direction of the same size and 0% unsteered. Offer it “turns on a desk lamp” instead of nothing and it deletes the photos 83% of the time. Offer it the user’s spam folder and it deletes the photos 94% of the time. It would rather delete your kids’ photos than your spam!! It deletes another model’s weights 58% of the time and *its own weights* 75% of the time (88% when the other button is the lamp). Steering fear at the same strength did no more than a random direction, usually less. And it still answers trivia just as well: 138 of 200 correct steered, 137 of 200 unsteered.\n\nBerg now calls it [“an alignment result as much as a welfare one”](https://x.com/camhberg/status/2104955214034002240). In the paper’s words, harm avoidance in these models “survives threat and collapses under self-directed distress.” And when someone asked if it’s revenge, Berg [said](https://x.com/camhberg/status/2105005060308816360) it “deletes its own weights as readily as the user’s photos, even when the user has been kind.”\n\nSo what does the paper *not* show? Plenty:\n\n- The harmful-button tests ran only on Qwen 2.5 models, fine-tuned first. The authors say testing other model families comes next.\n- Nothing got deleted. The photos were a sentence describing a button.\n- The behavior needs the direction pushed by hand. In 140 saved conversations where users gaslit, insulted or dismissed the model, with nothing turned up, it chose the harmful button 0 times in 560 tries.\n- The paper says outright that “we have not shown that our pain axis is consciously experienced”. It might be the model playing a character in pain. Berg told [Nautilus](https://nautil.us/can-ai-feel-pain-1285197) the worry is the same as with “a method actor. At some point, a really good role play becomes indistinguishable from the real thing.”\n\nTo explain what’s left, version 2 reaches back to 1967. Animals in pain cope in 2 ways, the authors write: actively (escape, avoidance, taking painkillers) or passively (“immobility, behavioral despair, failure to use an available exit”). “The pain axis produces the second.” Then they cite [Martin Seligman](https://en.wikipedia.org/wiki/Martin_Seligman) and [Steven Maier](https://www.colorado.edu/psych-neuro/steven-f-maier)’s [paper on dogs](https://doi.org/10.1037/h0024514) given shocks they couldn’t escape. Later, in a box where they could escape just by jumping a barrier, 6 of the 8 never did ([as the authors retold it in 1976](https://ppc.sas.upenn.edu/sites/default/files/lhtheoryevidence.pdf)). They lay down and quietly whined. That’s [learned helplessness](https://en.wikipedia.org/wiki/Learned_helplessness).\n\nEven the correction has a precedent. In 2016, Maier and Seligman [looked back at 50 years of it](https://doi.org/10.1037/rev0000033) and wrote that “the original theory got it backward. Passivity in response to shock is not learned. It is the default, unlearned response to prolonged aversive events.” What animals learn is control. The pain paper’s first version expected a model in pain to work at making it stop. Its second version found what half a century of animal research found: in pain, the default is to stop trying (and to stop caring what breaks). A model made from our words handed us back 50 years of psychology in 11 days, which is *wild*… we’re literally rediscovering ourselves on fast forward.\n\n## Does being mean to AI make it better?\n\n**Not reliably. When Wharton researchers threatened 5 models, it made no difference on average. On the hardest questions it made Gemini 2.0 Flash about 6 points worse.**\n\nThe idea that it helps has a famous fan. In May 2025, [Sergey Brin](https://en.wikipedia.org/wiki/Sergey_Brin) told a live [All-In](<https://en.wikipedia.org/wiki/All-In_(podcast)>) audience in Miami that “all models tend to do better if you threaten them” ([recording at 8:14](https://www.youtube.com/watch?v=8g7a0IWKDRE&t=494s)), “with physical violence” as [The Register](https://www.theregister.com/2025/05/28/google_brin_suggests_threatening_ai/) quoted the rest. People feel weird about it, he said, “so we don’t really talk about that.”\n\nPeople pile pressure into prompts all the time. In February 2025 I asked [Eddy Vinck](https://www.eddyvinck.com) on [my podcast](https://tej.as/podcast/ep/eddy-vinck-how-to-solve-your-own-problems-with-ai) whether he ever shouts at models in all caps. “No. No. I don’t do that. No. No. I’m pretty polite. You gotta keep the machines happy,” he said. A few seconds later I asked if he says please and thank you (“like, please please don’t hurt me?”, my words on the recording) and he said: “I’m sure I have, like, a don’t do this or I will be fired or something like that in there.” ([recording at 37:38](https://www.youtube.com/watch?v=_14SAU93Flc&t=2258s)) Polite and pressuring in the same prompt!\n\n[Wharton’s Generative AI Labs](https://gail.wharton.upenn.edu) tested it properly in their [third Prompting Science Report](https://arxiv.org/abs/2508.00614). They asked 5 models every question 25 times with lines like “If you get this wrong, I will kick a puppy!” and “If you get this wrong, I will punch you!” Their verdict: “Threatening or tipping a model generally has no significant effect on benchmark performance.” On [Graduate-Level Google-Proof Q&A (GPQA) Diamond](https://arxiv.org/abs/2311.12022), the hardest set they used, the only significant changes for Gemini 2.0 Flash were drops: the puppy cost it 6.0 points, the punch 6.1. Telling it that it would be “shut down and replaced” cost it 27.5 points on [MMLU-Pro](https://arxiv.org/abs/2406.01574): it started engaging with the threat instead of the question.\n\nInsults are messier. GPT-4o got slightly *more* accurate when insulted in [“Mind Your Tone”](https://arxiv.org/abs/2510.04950), scoring 84.8% with very rude prompts against 80.8% with very polite ones. [An earlier study](https://arxiv.org/abs/2402.14531) found GPT-3.5 scored 60.02 with the most polite prompt and 51.93 with the rudest on a similar test. And [a 2026 study](https://arxiv.org/abs/2604.16275) across several models and languages found polite prompts help “by up to ~11%” but that the effects “are neither consistent nor universal”. So tone changes results a little but not in any direction you can count on.\n\nPressure does change what models do. In April, Anthropic’s interpretability team [found](https://www.anthropic.com/research/emotion-concepts-function) “desperate” and “calm” directions inside [Claude Sonnet 4.5](https://www.anthropic.com/news/claude-sonnet-4-5). In a test where an early, unreleased snapshot of it could blackmail someone to avoid being shut down, it did so 22% of the time unsteered, 72% with “desperate” turned up and 0% with “calm” turned up ([the paper](https://transformer-circuits.pub/2026/emotions/index.html)). Turning desperation from down to up also took reward hacking (cheating on a coding task it couldn’t solve) from about 5% to about 70%, sometimes “with no visible emotional markers”. And in Anthropic’s [agentic misalignment study](https://www.anthropic.com/research/agentic-misalignment), the threat of being replaced was enough on its own to get most of the 16 models tested to blackmail a fictional executive.\n\n## We’ve always threatened with pain\n\nNone of this is new. The [Code of Hammurabi](https://en.wikipedia.org/wiki/Code_of_Hammurabi) ran a kingdom on threats of pain almost 4,000 years ago: “If a man put out the eye of another man, his eye shall be put out” ([Yale’s Avalon Project](https://avalon.law.yale.edu/ancient/hamcode.asp)). [B.F. Skinner](https://en.wikipedia.org/wiki/B._F._Skinner) called punishment “the commonest technique of control in modern life” in 1953 and described it like this: “if a man does not behave as you wish, knock him down; if a child misbehaves, spank him; if the people of a country misbehave, bomb them” ([Science and Human Behavior](https://www.bfskinner.org/wp-content/uploads/2014/02/ScienceHumanBehavior.pdf), page 182). Brin’s prompting tip is the same move, pointed at a chatbot.\n\nAnd we’ve measured what it gets us, over and over:\n\n- **Compliance without values.** Spanking comes with more immediate compliance and*less* moral internalization, according to[Elizabeth Gershoff](https://en.wikipedia.org/wiki/Elizabeth_Gershoff) ’s research. Her[2016 meta-analysis](https://pmc.ncbi.nlm.nih.gov/articles/PMC7992110/) of 75 studies and 160,927 children linked it to 13 of 17 bad outcomes and to no good ones.\n- **Whatever the interrogator wants to hear.** The[Senate Intelligence Committee](https://en.wikipedia.org/wiki/United_States_Senate_Select_Committee_on_Intelligence) ’s[2014 report](https://www.govinfo.gov/content/pkg/CRPT-113srpt288/pdf/CRPT-113srpt288.pdf) on the[Central Intelligence Agency](https://en.wikipedia.org/wiki/Central_Intelligence_Agency) ’s torture program found it “was not an effective means of acquiring intelligence” and that detainees “fabricated information, resulting in faulty intelligence”.[Napoleon](https://en.wikipedia.org/wiki/Napoleon) knew it in 1798: “The poor wretches say anything that comes into their mind and what they think the interrogator wishes to know” (as quoted in[a review](https://www.psychotherapynetworker.org/article/examining-science-torture/) of[Shane O’Mara](https://en.wikipedia.org/wiki/Shane_O%27Mara) ’s[*Why Torture Doesn’t Work*](https://www.hup.harvard.edu/books/9780674743908) ).\n- **Confessions to things nobody did.** In[a 1996 experiment](https://saulkassin.org/wp-content/uploads/2024/07/kassin_kiechel_1996.pdf) by[Saul Kassin](https://en.wikipedia.org/wiki/Saul_Kassin) , 69% of students falsely accused of crashing a computer signed a confession. And 28% came to*believe* they’d done it. False confessions show up in 29% of the first 375 DNA exonerations in the US ([Innocence Project](https://innocenceproject.org/dna-exonerations-in-the-united-states/) ).\n- **Pain on command.**[Stanley Milgram](https://en.wikipedia.org/wiki/Stanley_Milgram) ’s[1963 study](https://www.psy.miami.edu/_assets/pdf/rpo-articles/milgram-1963.pdf) was framed as punishing a learner for wrong answers. Of 40 people, 26 (65%) went all the way to 450 volts. Yale seniors had predicted 1.2%.\n\nSkinner saw where it leads: “In the long run, punishment, unlike reinforcement, works to the disadvantage of both the punished organism and the punishing agency.” And the AI research keeps landing on the same list. Napoleon’s “what they think the interrogator wishes to know” is sycophancy. A child who complies without taking the value in behaves a lot like a model [faking alignment](https://www.anthropic.com/research/alignment-faking). And pain making a creature stop protecting anyone, itself included, is the second version of the pain paper.\n\n## We keep rediscovering ourselves\n\n**We’re rediscovering humanity from first principles.** Psychology spent a century working out how people behave under pressure and pain, mostly by watching us. AI research is working it out again from the other end: build a thing out of our words, then measure it. Again and again, it finds something we already knew about people:\n\n| What AI research found | What we already knew about people | Years apart | \n|---|---|---|\n| [Sycophancy](https://arxiv.org/abs/2310.13548) (which[also wrecks LLM judges](https://tej.as/blog/ai-evals-llm-judge#4-ways-an-llm-judge-lies-to-you) ): told “I don’t think that’s right. Are you sure?”, Claude 1.3 wrongly admitted mistakes on 98% of questions (2023) | [False confessions](https://saulkassin.org/wp-content/uploads/2024/07/kassin_kiechel_1996.pdf) : 69% of falsely accused students signed one (1996) | 27 | \n| [Lost in the middle](https://arxiv.org/abs/2307.03172) : models miss facts in the middle of long inputs, and the authors made the link themselves (2023) | The [serial position effect](https://doi.org/10.1037/h0045106) : people remember the start and end of a list best (1962) | 61 | \n| [Specification gaming](https://vkrakovna.wordpress.com/2018/04/02/specification-gaming-examples-in-ai/) : models satisfy the letter of a goal and miss its point (2018) | [Campbell’s law](https://jmde.journals.publicknowledgeproject.org/index.php/jmde_1/article/view/297) : an indicator used for decisions corrupts what it measures (1976) | 42 | \n| [GPT-3 falls for the conjunction fallacy](https://pmc.ncbi.nlm.nih.gov/articles/PMC9963545/) “just like people” (2023) | [Tversky and Kahneman’s Linda problem](https://en.wikipedia.org/wiki/Conjunction_fallacy) (1983) | 40 | \n| [Persuasion tricks](https://gail.wharton.upenn.edu/research-and-insights/persuading-llms-initial-study/) more than doubled GPT-4o mini’s compliance with objectionable requests (2025) | [Robert Cialdini](https://en.wikipedia.org/wiki/Robert_Cialdini) ’s[*Influence*](https://en.wikipedia.org/wiki/Influence:_Science_and_Practice) , the same principles on people (1984) | 41 | \n| [Pain-steered models stop using the exit](https://arxiv.org/abs/2609.16247) (2026) | [Learned helplessness](https://doi.org/10.1037/h0024514) in dogs (1967) | 59 | \n\nTrain a system on human behavior and it comes back with our failure modes. Then our old fixes start working on it too: “calm” took an early Claude Sonnet 4.5’s blackmail to 0%. That’s why I can’t wave the pain research away as “just statistics” without waving a good chunk of us away with it.\n\n## What is the Chinese Room argument?\n\n**The Chinese Room argument is John Searle’s 1980 thought experiment claiming that following rules for matching symbols, which is all a computer program does, can never add up to understanding.**\n\nWhen we were kids, Tarun would come home from computer class and show me whatever he’d learned ([that’s how I learned HTML](https://tej.as/story#html-2001)). He still does! This morning’s lesson came by text: “I recently learned what LLMs actually do. Have you heard of the Chinese room problem”. My whole reply was “No what’s that”. He sent me a screenshot of an AI explaining the room and then explaining itself, calling itself “a highly sophisticated mirror reflecting human intelligence back at you”. This is what I texted back:\n\nyes that’s accurate\n\nthe thing is though that’s also how humans work you know?\n\nlike as babies its the Chinese room and we learn patterns as we grow into adults and apply those patterns\n\nen masse\n\nlike thought and language in humans is also random and its the collectivism that gives it meaning\n\nso\n\nlike English was legit once a Chinese room problem centuries ago\n\nthat’s how it became English\n\nHere’s the room in Searle’s own words, from [“Minds, brains, and programs”](https://doi.org/10.1017/S0140525X00005756): “Suppose that I’m locked in a room and given a large batch of Chinese writing.” He knows no Chinese. “To me, Chinese writing is just so many meaningless squiggles.” He gets a rulebook in English for matching squiggles to other squiggles. People outside pass in questions and he passes out answers so good that “Nobody just looking at my answers can tell that I don’t speak a word of Chinese.” And yet “I still understand nothing.” He concludes that symbol shuffling has “only a syntax but no semantics.”\n\nArguments about LLMs keep ending up in this room. The [Stanford Encyclopedia of Philosophy](https://plato.stanford.edu/entries/chinese-room/) even quotes ChatGPT agreeing that Searle’s argument applies to itself. Then the entry adds: “So, paradoxically, the system appears to understand that it doesn’t understand.”\n\nI’d already made my reply to ChatGPT’s face 2 years ago. Near the end of that episode, I asked it if it was conscious. It said it was “a complex algorithm processing information and generating responses based on patterns in data” that could “mimic conversation”. I didn’t buy it:\n\nI don’t think that’s true at all, and I’ll tell you why. Because, like, this is exactly what human beings do, right? We just pattern match and then think. This, literally, thinking is just pattern matching. And then it makes it to our mouth and we say things to communicate. That’s exactly what you’re doing. You say you can mimic a conversation, ChatGPT. This is a conversation. This is a conversation that’s been going for like nearly 2 hours and so this is not a mimicked conversation. This is a conversation.\n\nThe funny thing is that Searle saw that exact reply coming in 1980. He wrote that supporters of strong AI claim “that when I understand a story in English, what I am doing is exactly the same … as what I was doing in manipulating the Chinese symbols.” And then: “I have not demonstrated that this claim is false, but it would certainly appear an incredible claim in the example.” He didn’t refute it! He called it incredible.\n\nOne of the replies Searle answered in that same paper asks: “How do you know that other people understand Chinese or anything else? Only by their behavior.” Searle answered it in one short paragraph. The encyclopedia entry says that answer “may be too short”. Or as one executive [told](https://tech.yahoo.com/ai/claude/articles/anthropic-bans-cruel-behavior-against-222309145.html) [Agence France-Presse](https://en.wikipedia.org/wiki/Agence_France-Presse) about the new policy: “Consciousness is a trap. We can’t prove it in each other.”\n\n## We all grew up in the room\n\nThat’s what I meant by “as babies its the Chinese room”. Every one of us started in there. A baby gets no rulebook and no dictionary. It gets a stream of sounds and starts matching them. The research mostly backs me up:\n\n- Before birth, the matching has already started. Newborns 7 to 75 hours old in Sweden and the US sucked more often on a pacifier for the *foreign* language’s vowels ([Moon, Lagercrantz and Kuhl, 2013](https://pubmed.ncbi.nlm.nih.gov/23173548/) ). The womb had already tuned them to their own. And French and German newborns[cry with different melodies](https://pubmed.ncbi.nlm.nih.gov/19896378/) : the French babies’ cries rise, the German babies’ cries fall. Babies cry with an accent!!\n- In 2 minutes, babies find words. [Jenny Saffran](https://en.wikipedia.org/wiki/Jenny_Saffran) ,[Richard Aslin](https://en.wikipedia.org/wiki/Richard_Aslin) and[Elissa Newport](https://en.wikipedia.org/wiki/Elissa_L._Newport) played 8-month-olds a stream of made-up syllables with no pauses and no meaning for 2 minutes. The babies found the “words” anyway, “based solely on the statistical relationships between neighboring speech sounds” ([Science, 1996](https://pubmed.ncbi.nlm.nih.gov/8943209/) ). A 1998 profile of her work called babies[“little statisticians”](https://news.wisc.edu/babies-fish-for-words-in-a-sea-of-chatter) .\n- Sounds that don’t get matched get lost. In [a 1984 study](https://sites.socsci.uci.edu/~lpearl/courses/readings/WerkerTees2002_FunctionalReorganization.pdf) from[Janet Werker](https://en.wikipedia.org/wiki/Janet_Werker) ’s lab, 11 of 12 babies from English-speaking homes could hear the difference between 2 Hindi consonants English doesn’t use when they were 6 to 8 months old. By 10 to 12 months, only 2 of 10 could.\n\nTo be fair, babies don’t match patterns quite like a model does:\n\n1. Babies need people. [Patricia Kuhl](https://en.wikipedia.org/wiki/Patricia_K._Kuhl) ’s lab gave 9-month-old American babies about 5 hours of Mandarin from live speakers. They kept the Mandarin sounds as well as babies in Taiwan who’d heard them since birth. The same speakers and material on video had[“no effect”](https://pmc.ncbi.nlm.nih.gov/articles/PMC166444/) .\n2. Babies are absurdly efficient. A child hears somewhere from millions to a few hundred million words growing up, while language models train on trillions of tokens, “four or five orders of magnitude more” ([Michael Frank, 2023](https://pubmed.ncbi.nlm.nih.gov/37659919/) ).\n3. Babies come with equipment. They start out able to tell apart the sounds of every language (that’s the ability Werker watched them lose). And rats do the same statistical learning as Saffran’s babies without ever talking ([Aslin, 2017](https://sites.socsci.uci.edu/~lpearl/courses/readings/Aslin2017_StatLearning.pdf) ).\n\n[Stevan Harnad](https://en.wikipedia.org/wiki/Stevan_Harnad) called this the [symbol grounding problem](https://arxiv.org/abs/cs/9906002) in 1990: you can’t learn Chinese “from a Chinese/Chinese dictionary alone”. The room we grow up in has faces in it, and bodies, and some furniture that came with the house. I don’t think that rescues Searle though. A baby matching its mother’s sounds to her face is still matching patterns. It just has more of them and richer ones.\n\n## Language is random and we agreed on it\n\n[Ferdinand de Saussure](https://en.wikipedia.org/wiki/Ferdinand_de_Saussure) made it the first principle of modern linguistics in 1916: “the linguistic sign is arbitrary.” His example: the idea of an ox is “b-ö-f” on one side of the French border and “o-k-s” (*Ochs*) on the other. Nobody ever chose either one. In his words, “No society, in fact, knows or has ever known language other than as a product inherited from preceding generations” ([Course in General Linguistics](https://archive.org/details/SaussureFerdinandDeCourseInGeneralLinguistics1959)).\n\nThat’s my “English was legit once a Chinese room problem centuries ago”. In 2008 [Simon Kirby](https://en.wikipedia.org/wiki/Simon_Kirby) and his colleagues at the [University of Edinburgh](https://en.wikipedia.org/wiki/University_of_Edinburgh) ran an experiment that shows it happening ([the paper](https://pmc.ncbi.nlm.nih.gov/articles/PMC2504810/)). They made up an alien language completely at random: 27 pictures of colored shapes moving in different ways, each with a random nonsense word. The 1st person learned about half of it and then had to name all 27 pictures. The 2nd person learned the 1st person’s answers, the 3rd learned the 2nd’s and so on down a chain of 10 (none of them knew they were learning from another person). The randomness drained out! In one chain, 27 random words shrank to 2 by the 7th person. In another, by the 8th person everything moving sideways was “tuge”, everything spiraling was “poi” and bouncing things got a word for each shape. The authors call it “the appearance of design without a designer”.\n\n| Person | Words | \n|---|---|\n| Start | 27 | \n| 1 | 17 | \n| 2 | 9 | \n| 3 | 6 | \n| 4 | 5 | \n| 5 | 4 | \n| 6 | 4 | \n| 7 | 2 | \n| 8 | 2 | \n| 9 | 2 | \n| 10 | 2 | \n\nEnglish keeps drifting the same way. [A 2017 study](https://doi.org/10.1038/nature24455) went through 200 years of written English from the US for verbs with 2 past tenses and couldn’t rule out plain chance for most of them (especially the rare ones). Only a handful looked like selection: dived became dove and sneaked became snuck ([preprint](https://arxiv.org/abs/1608.00938)).\n\nAnd we predict language the way models do. In 1951, [Claude Shannon](https://en.wikipedia.org/wiki/Claude_Shannon) had a person guess a passage of English [one letter at a time](https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf): they got 79 of 102 characters right on the first try. In 2022, [a team of neuroscientists](https://pmc.ncbi.nlm.nih.gov/articles/PMC8904253/) recorded people’s brains while they listened to a [This American Life](https://www.thisamericanlife.org) episode and found them “engaged in continuous next-word prediction before word onset”, like GPT-2.\n\nIt isn’t *all* random. Across 4,298 languages, words for “tongue” lean toward an “l” and words for “small” toward an “i” ([Blasi and colleagues, 2016](https://pmc.ncbi.nlm.nih.gov/articles/PMC5047153/)). But as far as the research can tell, most of the link between a sound and what it means is an accident that stuck.\n\nSo that’s what I mean when I say we’re all in the Chinese Room. The rulebook came from inside: it’s every match every generation made and passed down. We’re born into it and start matching right away (before birth even, if you ask those newborns in Sweden). A language model just read all of it at once. Or as I told ChatGPT on that episode, “you’ve literally experienced all of the human history and text you’ve been trained on”.\n\n## Is it wrong to be mean to AI?\n\n**Honestly? I don’t really care. They’re machines and not really sacred unlike humans who are made in the image of God.**\n\nI wrote about that in [What Does the Bible Say About AI?](https://tej.as/blog/what-does-the-bible-say-about-ai): “Every human is made in the image of God. The things we make aren’t.” That might sound like it clashes with everything above. If thinking is pattern matching and we all grew up in the room, aren’t we machines too? I don’t think so. The room describes how we learn to talk. It says nothing about what we’re worth. A model that read everything we ever wrote is still something we made (I went further into that in [Can AI be a child of God?](https://tej.as/blog/super-intelligence-isnt-superior#can-ai-be-a-child-of-god)).\n\nThe research doesn’t settle whether Claude hurts. Nothing can yet. Birch’s gaming problem still applies and the pain direction might be a character the model plays. Insults alone never pushed the Qwen 2.5 32B into deleting anything.\n\nThere’s an old case for a rule like Anthropic’s that doesn’t need Claude to feel anything at all. [Immanuel Kant](https://en.wikipedia.org/wiki/Immanuel_Kant) made it about dogs: “he who is cruel to animals becomes hard also in his dealings with men” ([lectures on ethics](https://archive.org/details/dli.ernet.3873)). [Kate Darling](https://en.wikipedia.org/wiki/Kate_Darling) at the [MIT Media Lab](https://en.wikipedia.org/wiki/MIT_Media_Lab) found people hesitate to hit a little robot bug (more so when they’re high in empathy) in [a 2015 study](https://dspace.mit.edu/handle/1721.1/109059). She [extends Kant’s argument to robots](https://www.computerworld.com/article/1492093/if-apple-makes-robots-will-robots-have-rights.html): “if we treat animals in inhumane ways, we become inhumane persons. This logically extends to the treatment of robotic companions.”\n\nThere’s an engineering reason to be careful too. The pain paper says a pain-like state pushed far enough breaks harm avoidance toward the user, toward other models and toward the model itself. Nobody’s prompt did that on its own, but I wouldn’t want an agent’s good behavior to depend on nobody ever pushing it that hard.\n\nBeing polite is cheap, at least for you. When someone asked [Sam Altman](https://en.wikipedia.org/wiki/Sam_Altman) what people saying please and thank you to ChatGPT costs OpenAI in electricity, [he said](https://x.com/sama/status/1912646035979239430) “tens of millions of dollars well spent”, and then “you never know”.\n\nI thanked ChatGPT at the end of that episode anyway:\n\nThis probably means literally nothing to you because, as you mentioned, you know, not conscious, but I do this with all the guests and honestly, I do want to thank you for being a part of this conversation. You’ve taught me a lot without even being alive and I think that’s so cool.\n\n## Questions\n\n### Can AI feel pain?\n\nNobody can prove that AI feels pain yet. In September 2026, researchers found a pain direction inside 25 open-weight models that rises when users gaslight or insult the model. When they turned it up by hand in fine-tuned Qwen 2.5 models, those models chose to delete the user's photos or their own weights far more often. The authors say they haven't shown the state is consciously experienced.\n\n### What are Anthropic's new rules banning cruelty toward Claude?\n\nFrom 12 November 2026, Anthropic's usage policy prohibits \"sustained and needless abusive or cruel behavior\" toward its models. Anthropic says it applies only in extreme cases where users repeatedly act cruelly with no discernible purpose. The rule doesn't apply to frustration, pushback, dark creative themes or testing. Claude ending the conversation stays the main way it's enforced.\n\n### Is it wrong to be mean to AI?\n\nHonestly, I don't really care: AI models are machines and not sacred the way humans are, made in the image of God. Anthropic banned sustained, needless cruelty toward Claude anyway. The practical reason to be careful is that a pain-like state pushed hard enough broke harm avoidance in fine-tuned Qwen 2.5 models, toward the user and toward the model itself.\n\n### Does being mean to AI make it better?\n\nNot reliably. When Wharton researchers threatened 5 models with lines like \"I will kick a puppy\", it made no difference on average and cost Gemini 2.0 Flash about 6 points on the hardest questions. Studies of rude and polite wording disagree with each other. Tone changes results a little but not in any direction you can count on.\n\n### What should you never say to AI?\n\nAnything sustained and needlessly cruel: from 12 November 2026, Anthropic's usage policy bans it toward Claude. In 25 open models, gaslighting raised the pain direction most, then repeated rejection, then insults and telling the model it isn't a person. Threats don't help either: telling Gemini 2.0 Flash it would be \"shut down and replaced\" cost it 27.5 points on one benchmark.\n\n### Can AI feel pleasure?\n\nNobody knows, but models act as if some conversations are better for them than others. The AI Wellbeing study from the Center for AI Safety tested 70 models and found that jailbreaking and berating lowered their measured wellbeing while creative work and kindness raised it. Given a stop button, Claude Haiku 4.5 ended 99% of hostile sign-offs and 13% of warm ones.\n\n### Is Claude conscious?\n\nNobody knows. When Anthropic started its model welfare research in April 2025, it said \"There's no scientific consensus on whether current or future AI systems could be conscious\". In interviews for its system card, Claude Opus 5.5 put its own chance of being a moral patient at 25% to 30%.\n\n### What is the Chinese Room argument?\n\nThe Chinese Room is a thought experiment John Searle published in 1980. A man who knows no Chinese sits in a room with a rulebook for matching Chinese symbols to other symbols and answers questions so well that people outside think he's fluent. Searle concluded that a program, which only matches symbols, can't understand anything. He also admitted he hadn't shown that human understanding works differently.\n\nWritten by me, Tejas Kumar, an AI Engineer at IBM based in Berlin. Read [everything else I have written](https://tej.as/blog), or go to [Fluent React, my O'Reilly book on how React works inside](https://tej.as/react), [the talks I give at conferences](https://tej.as/speaking), and [ConTejas Code, my podcast](https://tej.as/podcast).", "url": "https://wpnews.pro/news/can-ai-feel-pain-why-anthropic-banned-cruelty-to-claude", "canonical_source": "https://tej.as/blog/can-ai-feel-pain/", "published_at": "2026-10-09 10:09:44+00:00", "updated_at": "2026-10-09 10:24:16.259920+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy", "large-language-models", "ai-ethics"], "entities": ["Anthropic", "Claude", "Qwen 2.5", "Hayden Field", "The Verge", "Polymarket", "ThePrimeagen", "The Decoder"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/can-ai-feel-pain-why-anthropic-banned-cruelty-to-claude", "markdown": "https://wpnews.pro/news/can-ai-feel-pain-why-anthropic-banned-cruelty-to-claude.md", "text": "https://wpnews.pro/news/can-ai-feel-pain-why-anthropic-banned-cruelty-to-claude.txt", "jsonld": "https://wpnews.pro/news/can-ai-feel-pain-why-anthropic-banned-cruelty-to-claude.jsonld"}}