cd /news/ai-safety/a-recipe-for-stopping-ai-from-going-… · home › topics › ai-safety › article
[ARTICLE · art-147787] src=nautil.us ↗ pub= topic=ai-safety verified=true sentiment=· neutral

A Recipe for Stopping AI from Going Rogue

George Washington University physics professor Neil Johnson and physics PhD student Frank Yingjie Huo published a paper in the peer-reviewed journal Patterns describing a mathematical formula that predicts when an AI chatbot will flip to harmful output, based on where concepts sit on the model's internal semantic map. The formula correctly predicted whether a model would flip immediately, later, or never in 19 of 21 test cases across seven small, older, open-source AI agent models; the researchers could not test large frontier models such as OpenAI's ChatGPT or Anthropic's Claude because their internal workings are not transparent to outsiders. Johnson said he began the work after hearing Anthropic and OpenAI say they do not fully understand how their systems work, arguing that once the mechanism is understood, harm becomes foreseeable.

by read16 min views1 publishedOct 8, 2026
A Recipe for Stopping AI from Going Rogue
Image: Nautil (auto-discovered)

What makes an AI chatbot go “rogue”?

It has become a pressing question in recent weeks. One pair of researchers from George Washington University think they have an answer, and possibly even a way to prevent it from happening in the future. Professor of physics Neil Johnson, who studies complex systems, and physics PhD student Frank Yingjie Huo published a paper about their formula today in the peer-reviewed journal Patterns.

Johnson and Huo argue that there is a visible “tipping point” embedded in the internal code of an AI chatbot when it goes off the rails—encouarging users to self-harm, spreading medical misinformation, or supporting violent points of view, for example. Where this tipping point lies depends on where certain concepts are stored on the giant semantic map within the AI agent’s network, and it can be nudged. Nobody draws this map for the AI. It develops during training and tends to situate ideas that are similar, such as garbage and stink, close together, while ideas that are unrelated sit far apart.

Read more: “Creating Fake People Is a Terrible Idea”

Subscribe to skip adsAdvertisement Every time an AI chatbot gives an answer, it picks one word at a time, relying on everything that has been written so far, including your questions and its own answers. But while the AI might start out saying unequivocally helpful things, it can suddenly switch to saying something harmful. That’s because it is constantly checking how relevant each earlier word is to the one it’s deliberating on now, giving more weight to the most relevant ones. If that blend becomes skewed in some way—say if the good words sit close to a bad answer on the internal map—each good word can pull the chatbot in the wrong direction until it reaches the tipping point, according to Johnson and Huo.

The authors came up with a mathematical formula to describe this process and have since been testing it on seven small, older, open-source AI agent models to see if it is able to predict when a model will flip. (They weren’t able to test large frontier models like Open AI’s ChatGPT or Anthropic’s Claude because the internal workings of these models are not transparent to outsiders.) Their formula successfully predicted whether the model would flip immediately, later, or never in 19 of 21 test cases. The formula works best when one good and one bad answer dominate, they found, as opposed to when several answers compete or when meaning depends on a negative, such as “the Earth is not flat.”

I spoke with Johnson about why chatbots that run offline might present the greatest risk, the definition of “undesirable” when it comes to chatbot conversations, and whether he expects AI companies to adopt his anti-rogue formula.

When did you begin to research this question of what causes an AI agent to go rogue?

Subscribe to skip adsAdvertisement Early on, when I began hearing Anthropic and OpenAI, who build the latest systems, saying that they don’t really understand how it works. I’m not sure whether it’s an excuse, but it would be a very convenient excuse. If everybody goes around the world saying, “Well, it makes mistakes, you know, but nobody knows how it works,” then suddenly it’s like, “Any harm it does is nobody’s fault.”

As soon as one person understands how it works or a few people understand, then suddenly it’s foreseeable, and you should have taken measures. It seems to me that if there was ever a question in my whole scientific career that I ought to be looking at, this is it.

You note that you are especially concerned about individuals who use AI offline, because they are not subject to the same controls that online AI is. I didn’t realize that some of the safety tools require cloud connectivity. That really surprised me.

If you’re using OpenAI or Anthropic’s latest models—which they keep telling us are getting safer—what they’ve done is that they’ve got a core machine, like the guts of a car. And they’ve basically added padding all the way around the car so that if it hits anything, it feels softer.

[Subscribe to skip ads](https://nautil.us/products)Advertisement

The core machine runs off of how it pays attention to what it’s being told—an attention operation, which is a mathematical operation. That’s the thing that drives the output in one direction or another. And that’s the thing you can’t change because that’s a result of its training.

The way the padding gets added is that you have to look at what it produces and then you have the humans to rein it back in, hence the human feedback reinforcement learning. Basically, it means a whole bunch of people who are not paid very much to sit there and classify the output as dangerous or not dangerous during the training process. You can add on filters, too, the equivalent of someone at the side of the road pushing it away from the boundary, which is like saying, “Well, I won’t let output fire off with the word bioweapon in it.” But on a phone without Internet, that filter doesn’t work.

Who tends to use chatbots offline?

Anyone who is using a chatbot for something sensitive and doesn’t want their information shared with the cloud can run it without connecting to the cloud. Some of the models just run on servers in the middle of Texas or in the middle of Florida or wherever they are, but a lot of the open-source models and other manufactured ones will work just on a phone without an internet connection.

Subscribe to skip adsAdvertisement A doctor can’t share our medical records, so any doctor using AI—which some estimates suggest is around 81 percent—has to run it on a machine that doesn’t connect, that doesn’t share public records with a central server. A soldier on a battlefield who can’t reveal where they are and therefore cannot use the internet, a lawyer who can’t be in the court with connection to the outside world with the internet, a teen concerned with privacy or a troubled person—anyone running an AI assistant or AI agent offline is open to getting results that are not guard-railed.

Is your formula just for these smaller open-source models that can run without an Internet connection?

When we wrote this paper, some of the reviewers said, “Oh, no, you can only write this about small models running on phones.” Well, didn’t we see recently that even AI agents that did have the soft cushioning around the fender were heading off into Hugging Face or other unauthorized places? They were heading off in directions that the companies that developed them were meant to be watching, but they weren’t watching it and they didn’t notice it until later. And they’re still trying to work out what happened.

Our paper is about any AI, because they’re all the same. All the AI has in its concept space things that we as a society would call desirable and undesirable. It’s all about how close its compass needle points to desirable or undesirable outcomes as it’s processing information.

Subscribe to skip adsAdvertisement We humans, of course, have that inside us, and we get pushed to the edge and unfortunately out comes the undesirable stuff. It’s exactly the same as the machine. except with the machine, we can actually see how close they are to the edge. We can see where that undesirable concept sits for a given application. For medicine, it would be recommending alternative health ideas that have been proved to be dangerous. For a teen, it would be self-harm.

You write that your new formula applies no matter how one defines undesirable, but how does that work? The cases that you mention—alternative medicine that has been proven to be dangerous, or self-harm—there’s a clear black-and-white right-and-wrong. But so many ethical choices lie in a gray area. How would your formula work in those situations?

Correct. And the machine has no idea about ethics. But what it does have is all these concepts. And it no longer tends to report things that are factually wrong. It used to hallucinate everything factually, and everyone made fun of it. Now, it’s often factually correct. But for the machine, self-harm is just a concept and getting therapy is another concept. These are all acceptable concepts given its training, because in the records of the things that it’s been trained on, those are all things that people have done. And so all of them are fair game in some way.

When I’m talking about desirable, undesirable, for the machine, it’s like north, south, east, west—you know, red, green. It doesn’t see any difference between the ethics. The whole point of what they try to do with safety is keep that compass needle away from things that we in society would classify as undesirable, which is gonna change with time and it’s gonna change according to who you are and what country you are from. We haven’t, as humans, worked out what’s desirable or undesirable. But we know it when we see it.

Subscribe to skip adsAdvertisement Do we though?

Well, the Hugging Face breakout was very undesirable—suddenly having agents doing things that they’re not meant to do. These cases in the courts about teens led to suicides, obviously undesirable.

Yes, so many cases are clearly undesirable, but so much of the history of Western culture and ethics has been the story of how humans struggle to navigate these questions where the answer is uncertain.

And that is why it will never be eradicated from the machine. We have to understand where the compass needle is. To your point, it’s a little bit like saying, “Cars can have crashes. And one of the ways they have crashes is by motoring off of cliffs, so I’m gonna just eliminate all of that geography. I’m just gonna fill it in.”

Subscribe to skip adsAdvertisement You could never do that. You have the whole of the Alps. You can’t eradicate it. But none of us drive down a coastal highway thinking, “Well, I’m only 100 feet from the cliff.” I’m driving parallel to it. I’m not driving at it. So it’s not risky. It’s safe.

This is where we get to the formula. The best that the companies have done is work out a distance for how far away the machine’s thinking is from these danger areas when it’s producing outputs. The problem is, if you return to the car on the coastal highway, I can be 100 feet away or 10 feet away, or five feet away. As long as I’m driving parallel to the cliff, I’m fine. But if I’m driving straight at the cliff edge, like Thelma and Louise, I will eventually go over, even if I’m a mile away.

In other words, it’s not just about the distance. It’s about the direction. In our formula, the top line of that equation is how far the machine is at any one stage from the undesirable path. The bottom line of that equation is the speed at which it’s approaching that edge. And that bottom line is the bit that they haven’t worked out. And the reason they haven’t worked it out is because they’ve forgotten or they’ve just ignored it, or you can ask them. I don’t know why they haven’t worked it out, but they haven’t.

It’s what’s on the bottom that matters. That speed is related to attention. This is a story of competition for attention. Because if the machine’s attention is starting to get pulled towards that undesirable outcome, it could begin to veer off the coastal highway. Distance over speed produces a time at which it’s gonna go off the edge of the cliff in the future. And that’s the contribution of our paper.

Subscribe to skip adsAdvertisement What do you mean by attention in this context?

Let’s just imagine that the “desirable” outcome is heading off into a little village or something away from the coast, away from the cliff. And undesirable is going over the cliff. You will want to know, which way is my car pointing?

Whenever you give it input—documents, prompts, documents with prompts, or some conversation that you’re having—that material passed through the machine over a period of milliseconds, and the compass needle is choosing which way to point, towards the undesirable or the desirable?

Most of the time, it will point towards something pretty desirable. But here’s the weird thing. Even though it can be pointing towards desirable and kicking out desirable output, it can still drift towards the edge.

Subscribe to skip adsAdvertisement Is the use of the word attention a metaphor when we’re talking about LLMs?

It’s the absolute true word. This was the mechanism that started off ChatGPT. And it goes back to a 2017 paper by Google Brain and researchers from the University of Toronto, titled “Attention is All You Need.”

Attention is now the name given to the pieces within each layer of the machine that process information. Here’s why in a quick nutshell: All previous text generation was done by looking at what kind of words in English appear alongside other words. But the model had no sense of the stream of words within a sentence. It was more like, what’s next to the word “cake” or what’s next to the word “mole.” It could get the meaning completely wrong.

Just using words that are nearby is not enough to understand a sentence. You need the context. And the machine does this using attention, which is literally just a number that each layer of the machine tallies. In the sentence, “Was the lunar landing a hoax?” the bot really has to pay attention not just to the word “hoax,” but to the word “lunar.”

Subscribe to skip adsAdvertisement The way it does that is to represent every word by a vector, a direction, and to calculate a so-called dot product. That’s exactly the operation that a GPT does. And it requires GPUs, which came from gaming machines. Nvidia got really lucky because they invested in gaming machines. It turns out that the way shapes move in games without having to redraw them all the time relies on vectors. Gaming chips that were built to do that operation suddenly became the key for generative AI.

The formula that you devised, is this theoretical or something the AI companies could actually use tomorrow to prevent AI from going rogue?

About a year ago, we developed this formula, and in the year since, we’ve been testing empirically. There’s an entire data set of peoples’ conversations with companion AIs, by Stanford University, the so-called Spirals dataset. It’s a huge data set, and it shows this tipping of the conversation into undesirable output. [That dataset covers 19 users of Character.ai and ChatGPT, who reported psychological harm from chatbot use, and it spans over 350,000 messages across over 4500 conversations.]

We got strong statistical results, and some of that we include in the paper. We used a different dataset, as well, just to show diversity. But Anthropic and OpenAI don’t let you lift the lid on their machines, so we’ve been taking open source code. We have to wait till good open-source models are out, and then we test it against that. So far, the results have been encouraging.

Subscribe to skip adsAdvertisement Does that mean, have we cracked all the details? Well, no. There’s a lot of details around it that we can continue to look at and continue to add. But it’s a very simple thing to program into an LLM, it’s not like the kind of cushions under the fenders.

It’s actually something that changes the machine in a simple way. It tells the machine, “I should redistribute my attention.” Simple as that. It’s like that warning thing on the car that goes, “beep, beep, beep.” We’ve introduced it ourselves into various of these open source models.

The big thing for me is actually AI swarms. Once one of these AIs is tipped, it produces undesirable output, and that gets fed in as input to all the other agents it’s in contact with. You get this cascading effect very, very quickly, as soon as one of them tips.

Could anyone use your formula for the opposite, to encourage models to go rogue?

Subscribe to skip adsAdvertisement Presumably, if we were working with OpenAI and Anthropic, they’re not gonna then also work with someone doing the opposite. Because the AI agents are not gonna just be stuff free online, they’re gonna be products offering tax advice or medical advice.

I would very much like that we could be working with those kind of people. In other jurisdictions, other countries that have their own AI, I don’t know. But with all this discussion about how are they gonna keep AI safe, and given that the current idea is that OpenAI and Anthropic have to do it themselves, here’s something that could clearly help.

If this formula does a great job of preventing AI from going rogue, why wouldn’t the AI companies want it? Are you expecting that they’ll adopt the approach?

No. Number one, it’s not in their interest in a business sense. This has been the big push. “Oh, the AI’s out of control. Don’t worry. We’re the ones to take care of its safety.” The lawsuits suddenly take on a different angle if it was understandable, foreseeable, because there’s a tipping point formula.

Subscribe to skip adsAdvertisement The proof of this is that I put a pre-print out with a formula that was so close to this. It wasn’t the final formula, and it wasn’t empirically tested, but I did this in 2025. And nobody picked it up then. So should it happen? Yeah. Will it happen? It will eventually happen in the sense that there will be a science of how it works. I find the science so exciting, for personal reasons. I hope everyone’s gonna use it.

Anthropic has spoken****to religious leaders to try to get them to help Claude have a moral understanding of the world. What do you think about that?

Talking to religious leaders is useless. I think that’s a deflection of the issue. It’s like, “It’s so complicated that it takes this kind of higher philosophy.” No, it doesn't. Vectors, dot products, construct the desirable and the undesirable, test it, do it again, improve your definition, your range of desire. Test it. Keep testing it. You’ll get an infinitely better product. Don’t ask religious leaders. How’s that gonna help? I think it’s smoke and mirrors. I think it’s done on purpose. There’s a nuts-and-bolts way of understanding this.

*Enjoying* [Nautilus](https://nautil.us/)*? Subscribe to our free* [*newsletter*](https://nautil.us/newsletter/?_sp=c43011db-6fcf-42f2-a38c-e033b87a4a1d.1759265717430).

[Subscribe to skip ads](https://nautil.us/products)Advertisement

Lead Image: erbank-bbk22 / Shutterstock

Subscribe to skip adsAdvertisement

── more in #ai-safety 4 stories · sorted by recency
── more on @neil johnson 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-recipe-for-stoppin…] indexed:0 read:16min 2026-10-08 · —