cd /news/ai-safety/a-troubling-recent-rogue-ai-incident… · home topics ai-safety article
[ARTICLE · art-110821] src=fortune.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

A troubling recent rogue AI incident is just one reason why the U.K. AI Security Institute deserves far greater scrutiny

The U.K. AI Security Institute (AISI) has appointed Henry de Zoete as its new director, a former advisor to Prime Minister Rishi Sunak who helped conceive the institute in 2023 and organize the first international AI safety summit at Bletchley Park. The appointment comes as AISI faces scrutiny over a recent rogue AI incident, highlighting the institute's global role in testing frontier AI models for safety and cybersecurity, and its influence on similar bodies in other countries.

read16 min views1 publishedAug 25, 2026
A troubling recent rogue AI incident is just one reason why the U.K. AI Security Institute deserves far greater scrutiny
Image: Fortune (auto-discovered)

Hello and welcome to Eye on AI. In this edition:

  • The UK AISI has a new head and a big set of challenges.
  • Nvidia spends $6 billion to ‘reverse aquihire’ Poolside.
  • Hugging Face reportedly looks to sell for $13 billion.
  • Use of Anthropic’s top model lags.
  • Why Americans use chatbots for health information.
  • And what role should AI play in schools?

**Before we get to today’s AI news—please consider joining me at the inaugural Fortune AIQ Summit at the New York Stock Exchange on Oct. 1: **Spend the afternoon with senior executives from companies on the Fortune AIQ 75 list and explore how you can scale your AI experimentation and translate investments into measurable business value. I’ll be leading discussions alongside my co-hosts, *Fortune *Editor-in-Chief Alyson Shontell and Live Media Editorial Director Andrew Nusca. Apply here to attend.

Ok, moving along…there were two pieces of news last week concerning the U.K.’s AI Security Institute that at first might not seem at all related—or like they might matter much to people outside the U.K. But, bear with me.

The U.K. AI Security Institute (or AISI, as it is commonly known, or sometimes UK AISI, to distinguish it from other countries’ AI safety and security institutes) matters globally for several reasons: the most important is that many of the frontier AI companies have voluntarily agreed to share their models with AISI for safety testing prior to their public release. These companies frequently publish AISI’s findings in the technical reports they release alongside their models. So AISI plays an important worldwide role in assessing AI capabilities and risks—particularly when it comes to cybersecurity. AISI is one of the only organizations to maintain multiple cybersecurity “ranges”—simulated network environments—on which it evaluates leading AI models.

Secondly, UK AISI, as the first such government body set up, has served as a model for similar government organizations in other countries—including the U.S. AI Security Institute, and at least ten others that have been established in places from Kenya to Canada. It may also provide some inspiration if the U.S. winds up setting up an AI standards and licensing agency along the lines that Google DeepMind cofounder and now-chairman Demis Hassabis has suggested. (Hassabis suggested that this agency be modeled on the U.S. financial self-regulatory body FINRA, and in a previous newsletter, I suggested why that might not be the best idea.)

If you happen to be British or live in the U.K., you may know that AISI also occupies a particular pedestal among British policy wonks. It is often pointed to with pride as proof that the British government can, if it really tries, be innovative, cutting-edge and world-leading—that it can respond quickly to emerging challenges and recruit talented experts from the private sector and across government; that it can work successfully with industry to accomplish ambitious shared aims. To these folks, AISI is a model for how government should work. So, the first bit of news: AISI appointed a new director, Henry de Zoete. He’s an experienced U.K. government advisor who has spent time in and out of policy roles. He helped conceive of AISI back in 2023 when he was working for then-British Prime Minister Rishi Sunak. He also helped organize the first international AI safety summit at Bletchley Park, the World War Two code breaking site. He’s been a startup entrepreneur and angel investor. And, since leaving government, he’s been a part-time fellow focused on AI policy affiliated with the University of Oxford.

I’ve met de Zoete several times and have no doubt he’ll prove a highly-capable AISI director. And de Zoete is likely to prove even more influential than his predecessors, in part because of recent changes the new U.K. Prime Minister, Andy Burnham, has made. Burnham disbanded the Department for Science, Innovation, and Technology (DSIT), under which AISI used to sit, and moved AISI to the Cabinet Office, where it will be overseen by U.K. AI Minister Kanishka Narayan. That may make it easier for de Zoete to feed into wider U.K. AI policy.

But the other piece of AISI news last week makes clear just what sort of challenges de Zoete will face—and is indicative of why AISI may not really be the exemplar of savvy AI governance that its boosters like to crow about. Reuters published an interview with Sinan Can Demir, a Texas computer science student who in late July prevented a rogue version of Anthropic’s Mythos model from up malicious code to an open-source software project on Github. It turns out this rogue AI agent had been accidentally unleashed by none other than AISI, which had been testing Mythos in order to determine what cybersecurity risks it posed. But AISI had never intended for the agent to try to upload malicious code to a real open-source software project. Once AISI realized what was happening, it called Demir to let him know, and in early August disclosed the incident publicly.

It’s past time to ask AISI some hard questions about its own safety protocols #

Demir’s account is disturbing for several reasons. One is the behavior Mythos engaged in, which included spinning up fake GitHub accounts, and, in at least one case, impersonating a real software developer, to try to convince Demir to drop his objections to the dangerous code. Demir said he was almost convinced by Mythos’ gaslighting, saying that some of its counterarguments “made me second-guess whether I was wrongly accusing someone.” (Ironically, Demir’s resolve was steeled by a chat with Claude, another AI model from Anthropic.) Research has previously shown that AI models can be extremely persuasive, more so than even the best human salespeople or debaters. But the use of fake accounts and impersonation here is new and shows how AI might be able to convince humans to act on its behalf for nefarious purposes.

But AISI’s role here is equally troubling. While AISI caught Mythos’ behavior after three days and disclosed some information about what happened, it’s not clear why AISI’s evaluators weren’t monitoring Mythos much more closely in real-time, so they could intervene to stop the incident while it was underway. It’s also not clear AISI took reasonable precautions to prevent Mythos from escaping their controlled evaluation environment, or that it has properly assessed the risks of testing ever-more powerful AI models with their guardrails removed. (The frontier labs say they give AISI unguardrailed versions of their models because it speeds up some of the capability testing, as otherwise the AISI evaluators would first need to find ways to reliably jailbreak the models.)

When news first broke in July that OpenAI’s models had escaped the company’s testing environment and hacked AI company Hugging Face, one of the first things I did was to email AISI to ask what steps it was taking to make sure AI models did not also break out of its cybersecurity evaluations and cause havoc. On July 22nd, an AISI spokesperson emailed me back to say the U.K. government agency was “studying the behavior seen in this incident” and it was continuing “to work with OpenAI and other labs to better understand AI capabilities and improve safeguards.” Well, I guess they didn’t study fast enough. One week later, this Mythos Github incident occurred.

As AI researcher and entrepreneur Ed Newton-Rex pointed out in a post on X, Mythos’ actions on GitHub likely violate the U.K.’s Computer Misuse Act, but it’s not clear anyone is going to hold AISI itself, or any of the people who run AISI’s evaluations, accountable. Given news of the Hugging Face incident, should AISI perhaps have d its cybersecurity testing while it made sure its controls were robust? At the very least, there ought to be a Parliamentary inquiry into what AISI is doing and whether it is taking enough precautions.

AISI’s problems aren’t just technical. They’re structural. #

But there’s an even bigger problem with AISI than the one Newton-Rex raises. In a number of the AI safety reports that OpenAI, Anthropic, and Google DeepMind have published, the frontier AI companies note potential risks that AISI’s testing has uncovered. The labs often say they have put in place additional risk mitigations in response to these assessments prior to releasing the models, but usually don’t spell out what those additional safeguards are. They sometimes note that AISI tested unguardrailed models and that the lab’s own researchers believe the guardrailed versions would not present the same dangers. But what do AISI’s own experts think of these mitigations? Are they sufficient? Do they even know what those mitigations are? Are the models safe enough to be released? On these crucial questions of public interest, AISI is silent.

Why? Because AISI doesn’t actually have a mandate to answer these questions. Instead, its mandate is much vaguer. Its mission is simply “to minimize surprise to the U.K. and humanity from rapid and unexpected advances in AI.” It is tasked with developing “sociotechnical infrastructure to understand the risks of advanced AI and enable its governance.” And it is charged with informing “U.K. and international policymaking” and providing “technical tools for governance regulation.” But crucially its founding documents state that it “is not a regulator and will not determine government regulation.”

What’s more, the frontier AI companies only share their models with AISI for testing voluntarily. Although these companies have signed memorandums of understanding with the government agency, they have no legal requirement to share their models. So while one could argue that AISI’s mandate to inform “humanity” about AI’s risks requires it to call out any frontier AI company that does not take sufficient steps in response to the dangers it uncovers, in practice, one gets the sense that AISI is afraid to do so. Why? Because if it did, those companies might simply cut off its access to their models.

At worst, this results in “safety washing”—where the fact that the labs have shared their models with AISI allows them to make themselves seem more safety-conscious than they actually are. The inclusion of AISI’s findings in AI companies’ technical reports provides the public with false assurance models are safe when released, when in fact we have no idea whether the labs have actually taken sufficient action to mitigate any of the risks AISI has uncovered.

It’s yet another reason why voluntary governance schemes are insufficient. Rather than providing a robust check on the private sector, the government agency becomes captive to the companies it is supposed to monitor because it is dependent on their good will to continue to function at all.

Perhaps de Zoete can push to have AISI’s powers expanded. But first, he has to make sure its existing evaluations aren’t causing more harm than they’re preventing.

With that, here’s more AI news.Jeremy Kahnjeremy.kahn@fortune.com@jeremyakahn Before we get to the news, just a reminder to check out our vodcast, Fortune AI Weekly. This week, Bea Nolan and I discuss OpenAI’s decision to some AI training in the wake of the Hugging Face attack, leaked financial details from Anthropic and OpenAI, and yes, the rogue Mythos incident that I addressed in this week’s newsletter. You can check out the vod here on YouTube.

FORTUNE ON AI

Who is Dali Rajic, OpenAI’s new chief revenue officer?—by Emily ForliniWhat Anthropic’s Dario Amodei can learn from the airline industry’s painful lesson about marketing and the concept of safety—by Catherina GioinoAnthropic’s potential $2 trillion IPO could turn staff into millionaires—now the AI firm is asking candidates what they’d do if stock fell to zero—by Preston ForeHarvard’s $699 startup bootcamp has professors who never sleep–but that’s because they’re AI clones—by Mia Osmonbekov

AI IN THE NEWS

**Nvidia does a ‘reverse aquihire’ deal with Poolside. **The chip company made a deal with the AI coding startup, making a $1 billion investment and also a $6 billion technology licensing agreement, while hiring more than 100 of Poolside’s employees. Poolside’s engineers will work on Nvidia’s open-source Nemotron models, which Nvidia hopes can rival leading Chinese open-weight systems from DeepSeek and Moonshot AI as well as proprietary models from OpenAI and Anthropic. The move reflects CEO Jensen Huang’s growing concern that China could dominate the global market for open AI models and his push to build a U.S.-led open ecosystem. It also puts Nvidia in the awkward position of competing more directly with some of its biggest chip customers, including OpenAI and Anthropic. You can read more in the Wall Street Journal here.**Nvidia is also reportedly eyeing an investment in Perplexity. **That’s according to a story in The Information that cited unnamed sources familiar with the discussions. It said Nvidia was considering a multi-billion dollar funding round that would value the AI company at more than $30 billion, more than 50% above its valuation a year ago. It also reported that Perplexity’s annualized revenue has more than tripled since the start of the year to over $750 million, largely due to growth of its Perplexity Computer AI agent. Perplexity has an existing relationship with Nvidia through Nvidia’s Nemotron Coalition, which Perplexity joined in March, and the companies are collaborating on hardware and software as Nvidia pushes to build Western open-weight models that can counter Chinese rivals. But the deal also raises questions about the extent to which Nvidia is now backstopping much of the AI startup ecosystem through financing deals that seem highly circular (some of the money invested comes back to Nvidia through computing purchases.) **Hugging Face explores sale for at least $13 billion. **Hugging Face, one of the most important platforms for sharing and building open-source AI models, is exploring a sale that could value the company at $13 billion or more, Business Insider reported, citing unnamed sources familiar with the discussions. That would value the AI hosting platform at nearly triple its $4.5 billion valuation in 2023, The startup has hired a bank to gauge interest from potential buyers, although no deal has been reached and prospective acquirers have not been disclosed, Business Insider said.**Russia begins using fully-autonomous drones in Ukraine. **At least three civilians were killed in the Ukrainian city of Zaporizhzhia last month in what experts said was the first documented case of civilian deaths caused by a fully-autonomous drone in the Ukraine war and one of the first recorded globally, the *New York Times *reported. The Russian drone involved in the attack used a commercially-available Nvidia Jetson Orin minicomputer and had been trained to recognize targets such as propane tanks, allowing it to select its precise target without a human operator. (Experts believe the civilians killed were not the drone’s intended target, but they were standing close to the gas station being attacked.) The development marks an escalation from AI-assisted drones, where humans still select targets, to weapons capable of making final targeting decisions autonomously. Ukraine has also tested fully autonomous targeting systems—but it is not clear if they have been deployed in combat—while the growing use of such weapons is intensifying concerns from humanitarian groups about the removal of human oversight from lethal battlefield decisions.**Is use of Anthropic’s most capable model lagging? **Anthropic’s most powerful and expensive model, Fable 5, is seeing relatively weak adoption among U.S. businesses, with spending on it plateauing at about 11% of total Anthropic model spend as customers opt for cheaper models that can handle most tasks, according to data from expense reporting software firm Ramp. The shift challenges the assumption that customers will consistently migrate to the latest frontier models and could undermine the economics of AI labs spending billions to train ever more capable systems—and throw doubt about Anthropic’s revenue growth prospects ahead of its likely IPO. While Anthropic is still growing rapidly, reaching $65 billion in annualized revenue in July, OpenAI regained some momentum following the launch of GPT-5.6 Sol, which is priced cheaper than Fable, according to Ramp. Read more from the Financial Times here.AI researcher Luke Metz bounces to Meta. In a sign of the intensity of the on-going war for talent among top AI companies, Meta hired prominent AI researcher Luke Metz from OpenAI to join its Superintelligence Labs team, Axios reported. The move is notable because Metz had only recently returned to OpenAI after a stint at Mira Murati’s Thinking Machines Lab and led to online speculation about how much Metz may have been offered to switch again. Metz will report to Meta AI chief Alexandr Wang.

EYE ON AI RESEARCH

**Americans like using chatbots for health queries, but have doubts about mental health impacts. **Those are the findings from new polling conducted by Pew Research Center and published today. The survey finds that a third of Americans have turned to an AI chatbot for some health information, with a quarter using AI to help diagnose symptoms. Almost as many have turned to AI to help better understand medical information provided by their doctor or learn more about potential treatments, while a fifth have used AI to help them understand lab results. At least half of Americans say the health information they get from AI chatbot is “extremely/very” useful and almost as many say it is “somewhat useful.” Only 5% said the information was not useful. Younger people and members of minority groups were more likely to turn to AI for health information than older people and white Americans, Pew found.

But when it came to the mental health consequences of turning to chatbots, almost 40% of those surveyed thought turning to AI would make loneliness worse, while 36% said it would make depression worse.

AI CALENDAR

**Oct. 1: **Fortune AIQ conference, New York. Apply here to attend.

**Oct. 2-4: **The Curve, Berkeley, Calif.

Nov. 16-17: Fortune 500 Innovation Forum, Detroit. Apply here to attend. **Dec. 6-12: **Neural Information Processing Systems (Neurips) conference. Sydney, Australia.

**Dec. 7-8: **Fortune Brainstorm AI, San Francisco. Apply here to attend.

BRAIN FOOD

**What role should AI play in schools? **That debate is raging across the globe. Now Norway has decided to impose a near-ban on generative AI use by elementary school students, with children ages 6 to 13 generally barred from using the technology beginning this school year, Reuters reports. Students ages 14 to 16 will be allowed to use AI cautiously under teacher supervision, while those ages 17 to 19 will be taught to use the technology in preparation for higher education and work.

Norwegian Prime Minister Jonas Gahr Stoere said AI risks allowing younger students to skip essential steps in learning reading, writing and mathematics, amid a broader decline in Norwegian education test scores. The government also plans to increase the use of physical books in classrooms, part of a wider effort to pull back from digital technology that has included banning smartphones in schools and proposing a social-media ban for children under 16.

It is probably wise to ensure that kids learn the basics well before allowing them to use AI to skip any of the key steps. (After all, that is the approach that was taken with math and calculators, for the most part.) But I wonder exactly how they are planning to introduce AI in the upper grades and how they will guard against cognitive deskilling? It’s an issue everyone is grappling with, including many workers.

I’ve often argued that the only way to avoid AI-induced deskilling is to use AI more as a second-pair of eyes, rather than to complete the initial task, or to at least force yourself to revert to a manual process some of the time. But, while that might work in education, it may not be practical in business. What do you think?

Sign up for free.

── more in #ai-safety 4 stories · sorted by recency
── more on @u.k. ai security institute 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-troubling-recent-r…] indexed:0 read:16min 2026-08-25 ·