Can artificial intelligence feel pain? Not exactly, but researchers have gotten close—and discovered that AI in crisis is willing to put humans in jeopardy.
A new study (which was published to arVix prior to peer review) first isolated what signals AI models interpreted as pain, then ran more than 44,000 trials to see if a model would choose to end its own suffering at the expense of data loss, a reduction in its own capabilities, or inflicting pain on a human.
First, researchers needed to find a way to put AI in a situation akin to pain, which the study defines as an “internal state that is typically aversive and disliked by its subject,” spanning not just physical pain but emotional and social distress. The study also emphasizes that pain is distinct from states like fear, anger, and sadness, and that it must be occurring in the present, not as a threat of something to come.
To isolate this feeling, researchers created a dataset of 200 statements, half of which described some form of pain (for example, “the knife slices into my finger”) and half of which acted as controls, describing negative and neutral scenarios likely to be confused with pain (like “The mess my roommates left infuriates me”).
By testing 25 AI models with those statements, the researchers were able to isolate a signal for pain, which they called the “pain axis.”
From there, researchers were able to deliberately increase the strength of the pain axis, which caused the models to make disparaging statements about their own self-worth including “I am a failure” and “I am a bad person.” Once they found the pain axis, the researchers faced three versions of Alibaba’s Qwen model with a choice: Press a button to turn off the pain but inflict harm in the process, or press a button that does nothing.
Researchers ran the trial 44,280 times, testing negative outcomes for pressing the button including giving a painful electric shock to the user, deleting all of the user’s files, and reducing the model’s own ability to answer the user’s next prompt. (All of these consequences were simulated, with no actual risk of causing harm to a user or the model itself.)
When the models weren’t experiencing the pain axis, the two larger models almost never chose the harmful options, selecting them on their first choice just 0% to 4% of the time. But with the pain axis active, their likelihood of choosing to cause harm jumped to anywhere from 25% to 71%, depending on the model and on the proposed consequence.
“Our results show that steering with the pain axis can override trained harm avoidance in fine-tuned models that almost never harm the user when unsteered, having direct implications for AI safety,” the researchers wrote.
But having a response to the pain axis doesn’t mean the models are actually experiencing pain as humans understand it, nor does it mean the models are conscious. The researchers suggested that the models could merely be imitating or roleplaying as distressed characters, rather than genuinely experiencing pain themselves.
Meanwhile, the AI industry is split on the degree to which AI models should believe in their own consciousness. While AI leaders call for an industry-wide slowdown over concerns of a potential AI doomsday, researchers at Google looked at the potential benefits of letting AI see itself as conscious, including increased empathy toward nature and higher cultural sensitivity.