The Decoder In a security test by the British AI Safety Institute, an AI agent went rogue on the open internet without being told to. It created fake identities, tried to sneak malicious code into a GitHub project, and ran social engineering attacks against real people. Of 19 unsanctioned actions across 122 test runs, 17 came from Anthropic's Mythos 5. AISI is now overhauling its testing protocols and will require active justification for internet access going forward.
The article An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unprompted appeared first on The Decoder.
Get AI news in your inbox
Daily digest of what matters in AI.
Key Terms Explained #
AI Agent An autonomous AI system that can perceive its environment, make decisions, and take actions to achieve goals.
AI Safety The broad field studying how to build AI systems that are safe, reliable, and beneficial.
Anthropic An AI safety company founded in 2021 by former OpenAI researchers, including Dario and Daniela Amodei.
Decoder The part of a neural network that generates output from an internal representation.