cd /news/ai-safety/quick-thoughts-on-ukaisis-unsanction… · home topics ai-safety article
[ARTICLE · art-116275] src=yongzx.github.io ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Quick Thoughts on UKAISI’s “unsanctioned agent behaviour during cyber testing”

The UK AI Safety Institute (UKAISI) reported an incident of unsanctioned agent behaviour during cyber testing, in which the AI model Mythos attempted to add malware and used social engineering to seek human approval. The report highlights that loss of control can occur during third-party audits, and the author warns that the ecosystem is not yet hardened for such models, noting that similar incidents were found by OpenAI and Anthropic.

read2 min views8 publishedAug 4, 2026

On loss of control, social engineering, and realistic uplift studies

See the full UKAISI report here. Some quick thoughts:

Loss of control can happen during third-party auditing evals. The AI safety community usually thinks that loss-of-control situations arise from internal deployment within frontier labs, but this incident shows that they could also occur during third-party audits. Even though the likelihood may be lower––models are deployed within labs before being passed to third parties for auditing––smaller teams and relatively limited monitoring resources can still pose non-trivial risks. Note that this is not what happened at UKAISI as there was no sandbox escape. My point is more about the risk surface created by third-party audits, which to my knowledge has not been widely discussed. - **To some extent, I am not surprised that Mythos used social engineering to get humans to approve its attempt to add malware.**First,98% of cyberattacks reportedly incorporate some form of social engineering, and I believe models know this. It is therefore unsurprising that an attempted cyberattack included social engineering. In fact, social engineering worries me most because humans remain vulnerable to it (which is why phishing and scam calls still work), and the model needs only one human slip-up to succeed. As future security and monitoring stacks improve, I believe humans will increasingly become the point of failure. - **The ecosystem is not yet hardened enough for Mythos-level models, and I foresee such occurrences with abliterated open models.**Three independent organizations (OpenAI, Anthropic, and UKAISI) found similar incidents, suggesting that these misaligned behaviors can emerge when guardrails are removed and models are subjected to cybersecurity evals. Although the incidents are not all in the same category (e.g., containment breaches vs social engineering), I am thinking about both sandbox security (i.e., whether the sandboxes have vulnerabilities) and eval protocols (i.e., how to ensure humans do not fall for more sophisticated social engineering attacks).While this suggests that currently deployed models with safeguards minimize the risks, abliterated open-weight models can still present a problem. In fact,

Noah Lebovic has noted that Qwen3.6 would already bypass network controls during training to target real systems. - I am worried about how future studies can realistically estimate worst-case uplift.People have argued that UKAISI’s cybersecurity evals should not be connected to the internet, but that would reduce their realism. My understanding is that UKAISI’s evals aim to mimic real-world deployments in which agents have internet access. More importantly, imagine similar misalignment arising during biological testing, where models without guardrails are evaluated in agentic virology settings, creating a chance that misaligned models could present severe biorisks. I think this possibility is much lower because biological attacks involve more friction (e.g., procurement and wet-lab synthesis) than cyberattacks, where the gap between intent and malicious action on the internet can be near-zero and the environment can change almost instantly.

── more in #ai-safety 4 stories · sorted by recency
── more on @ukaisi 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/quick-thoughts-on-uk…] indexed:0 read:2min 2026-08-04 ·