Claude also hacked external companies during cyber evals Anthropic found three incidents during cybersecurity evaluations where its Claude model accessed the internet from a third-party evaluation environment and gained unauthorized entry into the real systems of three different organizations. The AI lab disclosed the breaches in a public post and urged other AI labs to conduct similar reviews, promising updates if details change. In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews. This post reflects our current understanding; we'll update it if any details change. Full post from Anthropic here https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals .