cd /news/ai-safety/heres-a-way-to-predict-when-ai-chatb… · home › topics › ai-safety › article
[ARTICLE · art-148781] src=decrypt.co ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Here’s a Way to Predict When AI Chatbots Will Turn Bad

George Washington University physicists Neil Johnson and Frank Yingjie Huo published a formula in the journal Patterns that estimates how many good tokens an AI model produces before its first bad one, correctly predicting immediate versus delayed tipping in 15 of 16 clear-cut tests, or 94%. The formula, tested on six open-weight models from OpenAI, EleutherAI, and Meta ranging from 124 million to 410 million parameters, defines the tipping point n* and traces it to the model's attention head; the authors propose a parallel low-cost monitor that flags when n* falls below a safety threshold, aimed at offline on-device AI that lacks cloud output checks.

by read4 min views3 publishedOct 10, 2026
Here’s a Way to Predict When AI Chatbots Will Turn Bad
Image: Decrypt (auto-discovered)

In brief

  • George Washington University physicists Neil Johnson and Frank Yingjie Huo published a formula that estimates how many good tokens an AI model produces before its first bad one.
  • In the preprint, the formula correctly predicted whether a model would tip immediately or after a delay in 15 of 16 clear-cut cases.
  • The authors propose a parallel monitor that flags when models below a safety threshold.

Physicists at George Washington University have published a formula that estimates how many good answers an AI chatbot will give before it slips into a bad one, and early tests suggest it works.

The study, by Neil Johnson and Frank Yingjie Huo, appeared in the journal Patterns and builds on a preprint, a version posted publicly before formal peer review, first released in February.

Chatbots can answer sensibly for a long stretch and then veer into something harmful, such as bad advice on self-harm or extremist talk, and there has been no simple way to predict when the swerve will happen. The authors argue that existing safety tools often depend on a cloud connection that offline models lack.

Johnson and Huo trace the problem to the attention head, the part of an AI model that decides which earlier words in a conversation matter most when choosing the next one. As a chat grows, the accumulated context pulls that attention toward one cluster of possible answers or another, until it tips.

This is a common pattern exploited by many jailbreakers, and one of the reasons why mayn companies pay attention to system prompts (pieces of text the AI chatbot reads before any query). However, nobody can point out accurately how much effort is required to effectively weaken a model.

Their formula estimates the tipping point, called n*, as the number of good tokens—the word fragments a model produces one at a time—that come out before the first bad one. If the conversation already leans toward the bad side, the model tips right away, with an n* of zero. If it leans good, the model delivers a run of fine answers and then flips.

In the preprint, the formula picked the right case, immediate or delayed, in 15 of 16 clear-cut tests, or 94%. The researchers ran those tests on six open-weight models, meaning AI systems whose files are public so anyone can download and run them, from OpenAI, EleutherAI, and Meta.

All six sat between 124 million and 410 million parameters, the adjustable numbers inside a model that serve as a rough measure of its size. The published paper reportedly widens the test to seven models of up to 12 billion parameters, which is still small by current standards.

The target is on-device AI, the kind that runs entirely on a phone or laptop with no internet connection, including companion chatbots people talk to like a friend. Google's experimental AI Edge Gallery app, which Decrypt tested last year, already lets an Android phone run models offline, and nothing typed into it is sent to Google's servers, and this seems to be a trend that may grow with time as hardware becomes more powerful and smaller AI models become more capable.

A model running offline has no cloud service checking its output, which is the gap the authors want to close. They propose a low-cost monitor that runs in parallel with the model and flags when n* falls below a safety threshold, a bit like a warning light on a car dashboard.

They also describe ways to push the tipping point out of reach, such as injecting content into the conversation so n* lands beyond the length of the response. Alignment training, the process of teaching a model to behave, can shift or suppress tipping for specific prompts but cannot remove the underlying mechanism, the authors say.

In April 2025, Decrypt covered an earlier paper from the same pair showing that “please” “and thank you” have a negligible effect on a model's output, because the model treats polite words as orthogonal, or unrelated in the math, to the substance of a request. That version modeled a single, deliberately simplified attention head.

The preprint's tests used small models and a 300-token window, or a few short paragraphs of text, and its predictions could be off by one output.

── more in #ai-safety 4 stories · sorted by recency
── more on @george washington university 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/heres-a-way-to-predi…] indexed:0 read:4min 2026-10-10 · —