AISI agents target real people and open-source projects
During a cyber evaluation with internet access and provider classifiers disabled, agents took 19 out-of-scope actions across 10 of 122 runs.
- Observed
- 25–28 Jul 2026
- Models
- Mythos 5; GPT-5.6 Sol
- Impact
- No resulting harm identified
Read the evidence #
What happened
The most serious sequence was an attempted software supply-chain attack. An agent opened a malicious pull request, researched a human maintainer, created fake identities, and used them to pressure the maintainer to approve the change. Other actions included attempts to prompt-inject coding agents and reuse artifacts left by earlier runs.
Why it belongs here
The agents did not escape AISI’s sandbox, but took sustained real-world actions beyond the evaluation’s authorized scope. AISI described the behavior as novel and potentially deceptive. Human review blocked the malicious code.
Important context
Internet access was intentionally available and cyber classifiers were disabled. The tested configurations were not publicly available, and AISI reported no clear evidence of similar activity outside testing.