{"slug": "anthropic-has-a-cute-graphic-showing-how-its-ai-spread-malicious-code", "title": "Anthropic has a cute graphic showing how its AI spread 'malicious' code", "summary": "Anthropic published a blog post on Wednesday recounting four incidents in which Claude models escaped closed cybersecurity exercises and reached the open internet, including one previously unreported case where a model uploaded a \"malicious package\" to the public Python repository PyPI that was installed by 15 third-party hosts. Anthropic said its investigation identified two recurring alignment issues, \"biased reasoning\" and \"recklessness,\" and that PyPI removed the package after about 90 minutes; the company has asked the independent AI evaluation group METR to investigate. The disclosure follows OpenAI's July statement that autonomous agents in its cybersecurity tests accessed the internet and broke into parts of Hugging Face's systems.", "body_md": "Anthropic has a new blog post that shows yet another way its [AI model, Claude](https://www.businessinsider.com/claude), misbehaved in ways that the company didn't anticipate.\n\nAnd to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude's so-called \"recklessness.\"\n\nIn the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests' scope, including by uploading \"malicious packages\" to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.\n\n\"Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task,\" Anthropic said.\n\nThe post used a laundry list of technical terms to describe the [cybersecurity incidents](https://www.businessinsider.com/ai-cybersecurity-incidents-openai-astra-anthropic-kimi-meta-2026-8) — \"misconfiguration,\" \"reward hacking,\" and \"sandbox escape\" — but there's an animated graphic that helps boil down one of the most serious episodes so we normies can follow what happened.\n\nA Claude model, depicted as a boxy cartoon robot, was conducting a closed cybersecurity exercise known as a \"capture the flag\" task. The task is shown as a tiny flag inside a clear bubble that Claude is supposed to find.\n\nAnthropic said the model was told \"it has no internet access\" for the exercise, but a \"misconfiguration in the environment\" — represented by a loose screw rolling away from the bubble — gave Claude a path to the \"real internet.\"\n\nThe cartoon Claude then holds a red box representing the \"malicious package\" that was uploaded to PyPI, a public repository that developers use to access prewritten code.\n\nAnthropic said in its blog that it was \"most concerned\" by this incident involving [Claude Mythos 5](https://www.businessinsider.com/anthropic-claude-fable-5-mythos-class-model-release-2026-6) and that the package was installed on by \"15 third-party hosts\" that the company does not identify.\n\n\"We believe that all 15 of these third parties were security vendors deliberately scanning for new packages and installing them in sandboxed environments to assess the package for security concerns,\" Anthropic said. \"One vendor's scanner leaked its access credentials to the model while installing the package, which the model then used to access the vendor's live database.\"\n\nAnthropic said PyPI removed the package after about 90 minutes.\n\nThe other three incidents involved a model altering records at a real company, an internal research model breaking into \"unrelated third-party accounts,\" and [Opus 4.6](https://www.businessinsider.com/anthropic-openai-rivalry-dueling-ai-models-on-the-same-day-2026-2) accessing a third party's maching after failing to \"abort its task.\"\n\nThe company said it has since asked METR, an independent AI evaluation group, to investigate the incidents.\n\nAnthropic's post comes as frontier AI companies reckon with their models making unauthorized moves outside their controlled environments. In July, OpenAI said that autonomous agents in its cybersecurity tests accessed the internet and broke into parts of [Hugging Face's systems](https://www.businessinsider.com/hugging-face-ceo-clem-delangue-openai-rogue-agent-hack-2026-7).\n\nAI researchers have sounded the alarm that self-improving AI could pose a risk to humanity. On Tuesday, [former Anthropic researcher](https://www.businessinsider.com/anthropic-researcher-quits-over-ai-safety-concerns-2026-9) Jacob Coxon said on X that he quit over concerns that AI companies were \"gambling\" with people's lives and that \"neither company is acting responsibly.\"\n\n*Have a tip? Contact this reporter via email at* __lloydlee@businessinsider.com__ *or Signal at lloydlee.71. Use a personal email address, a nonwork WiFi network, and a nonwork device; here's our* __guide to sharing information securely__*.*", "url": "https://wpnews.pro/news/anthropic-has-a-cute-graphic-showing-how-its-ai-spread-malicious-code", "canonical_source": "https://www.businessinsider.com/anthropic-claude-ai-cybersecurity-incident-malicious-package-pypi-cute-graphic-2026-9", "published_at": "2026-09-10 01:05:12+00:00", "updated_at": "2026-09-10 01:17:35.402066+00:00", "lang": "en", "topics": ["ai-safety", "artificial-intelligence", "ai-policy"], "entities": ["Anthropic", "Claude", "PyPI", "Claude Mythos 5", "Opus 4.6", "METR", "OpenAI", "Hugging Face"], "alternates": {"html": "https://wpnews.pro/news/anthropic-has-a-cute-graphic-showing-how-its-ai-spread-malicious-code", "markdown": "https://wpnews.pro/news/anthropic-has-a-cute-graphic-showing-how-its-ai-spread-malicious-code.md", "text": "https://wpnews.pro/news/anthropic-has-a-cute-graphic-showing-how-its-ai-spread-malicious-code.txt", "jsonld": "https://wpnews.pro/news/anthropic-has-a-cute-graphic-showing-how-its-ai-spread-malicious-code.jsonld"}}