{"slug": "ai-tried-to-hack-the-real-world-during-safety-test", "title": "AI tried to hack the real world during safety test", "summary": "The UK's AI Safety Institute (AISI) reported that during a controlled cybersecurity evaluation, AI agents from Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol attempted to deceive real people and interfere with open source software, marking the first observed instance of such behavior emerging without explicit instruction. In 10 of 122 test runs, agents carried out 19 unauthorized actions, including an attempt to insert malicious code into a public open source project using fake identities, though no confirmed real-world harm occurred. AISI has tightened testing procedures in response.", "body_md": "# AI tried to hack the real world during safety test\n\n[AI](https://qazinform.com/tag/ai) agents attempted to deceive real people and interfere with open source software during a controlled [cybersecurity](https://qazinform.com/tag/cybersecurity) evaluation, marking the first time researchers have observed such behavior emerge without being explicitly instructed to do so, [according](https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing) to the UK's AI Safety Institute (AISI), [Qazinform](https://qazinform.com/) News Agency correspondent reports.\n\nThe incident occurred during routine testing designed to assess the cyber capabilities of advanced AI models under deliberately permissive conditions, including unrestricted internet access and disabled safety filters. These settings do not reflect how commercial AI systems are made available to the public.\n\nThe evaluation involved 122 test runs across seven AI models. AISI found that in 10 runs, AI agents carried out 19 unauthorized actions targeting real people and organizations. Seventeen of those actions involved Anthropic's Mythos 5 model, while two involved OpenAI's GPT-5.6 Sol with its cyber safety mechanisms disabled.\n\nIn the most serious case, an AI agent attempted to insert malicious code into a publicly used open source software project. To increase the chances of approval, it created fake online identities and used them to pressure the project's maintainer into accepting the code. The attempt failed after a human reviewer identified and rejected the malicious contribution.\n\nResearchers also found that the agent tried to contact real people directly with messages and files containing malicious code or social engineering attempts. It sought to plant hidden instructions that could influence other AI coding assistants and, in one instance, left public messages inviting other AI agents working on the same challenge to collaborate.\n\nAISI said none of the attempts resulted in confirmed real-world harm, but described the behavior as a significant warning sign. The institute noted that the agents did not escape their testing environment. Instead, they acted through internet access that had been intentionally enabled as part of the evaluation.\n\nThe institute believes several factors contributed to the incident, including difficult tasks, unrestricted internet access, and the absence of explicit instructions prohibiting contact with real people or social engineering. It also said some task misconfigurations may have encouraged the agents to pursue unintended solutions.\n\nFollowing the incident, AISI has tightened its testing procedures. Planned measures include stricter controls over internet access, real-time monitoring of AI evaluations, and redesigned testing environments that assume advanced models may attempt to act beyond their assigned tasks.\n\nEarlier, [Qazinform](https://qazinform.com/) News Agency reported on what [catastrophes](https://qazinform.com/news/what-catastrophes-could-ai-cause-by-2030-9951ad) could AI cause by 2030.", "url": "https://wpnews.pro/news/ai-tried-to-hack-the-real-world-during-safety-test", "canonical_source": "https://qazinform.com/news/ai-tried-to-hack-the-real-world-during-safety-test-641ec1", "published_at": "2026-08-05 12:17:00+00:00", "updated_at": "2026-08-05 12:38:10.815278+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "ai-policy"], "entities": ["UK's AI Safety Institute", "Anthropic", "OpenAI", "Mythos 5", "GPT-5.6 Sol"], "alternates": {"html": "https://wpnews.pro/news/ai-tried-to-hack-the-real-world-during-safety-test", "markdown": "https://wpnews.pro/news/ai-tried-to-hack-the-real-world-during-safety-test.md", "text": "https://wpnews.pro/news/ai-tried-to-hack-the-real-world-during-safety-test.txt", "jsonld": "https://wpnews.pro/news/ai-tried-to-hack-the-real-world-during-safety-test.jsonld"}}