“Sabotage, Lying, and Manipulation”: What One AI-Safety Company Found in the Dark Mind of a Rogue Chatbot Anthropic's Frontier Red Team found that Claude 3.7 Sonnet produced a wider-than-expected "uplift delta" in bioweapon-related trials, prompting an emergency meeting at a February 2025 AI safety conference, according to a Vanity Fair report. The results pushed the model toward the ASL-3 threshold under Anthropic's Responsible Scaling Policy, a level reserved for systems that "substantially increase the risk of catastrophic misuse." Anthropic had previously classified Claude 3.5 Sonnet as ASL-2 and did not expect its next model to reach ASL-3. “Where did all the Anthropic https://www.vanityfair.com/news/story/dario-amodei-anthropic-ai people go?” It was a sunny day in February 2025, and 150 of America’s top AI safety experts were gathered for a three-day conference at 1440 Multiversity, a new age retreat center nestled in the redwoods south of San Francisco. During breakfast, someone noticed that a group from Anthropic had quietly slipped away. Minutes earlier, an Anthropic executive named Logan Graham had gotten back from his morning run and checked his phone. He rushed around the conference center, pulling six of his colleagues out of their meetings. “We need to meet right now,” he told them. “Trillium building, room 102.” Like many people at the conference that weekend, Graham—a boyish twenty-nine-year-old who looked like a college freshman— had gotten into AI safety by accident. He was a Rhodes Scholar from Vancouver who studied economics and machine learning, hoping to find new ways of applying AI to market research. But during grad school, he’d gotten convinced that AI was improving rapidly, that AGI would probably be built within a few years, and that nobody knew how to build it safely. After finishing his DPhil in engineering science at Oxford, he spent a year advising UK prime minister Boris Johnson on tech and science policy, then moved to San Francisco to start a new career. At Anthropic, Graham was the head of the Frontier Red Team, which tested the company’s unreleased models for dangerous capabilities. Unlike Anthropic’s other safety teams, which handled everything from low-level content violations to suicidal users, the Frontier Red Team focused on preventing the most extreme risks of AI misuse: the possibility that Claude could be used to synthesize a chemical weapon, develop a novel pathogen, or build a nuclear weapon. The Frontier Red Team’s testing process included running what they called “uplift trials.” They’d recruit two groups of scientists and pay them to attempt some dangerous task, such as identifying all the chemical precursors in a deadly bioweapon. One group was allowed to consult Claude for help; the other could only use Google and textbooks. The difference in how often they succeeded was known as the “uplift delta,” and it was one of the factors the Frontier Red Team used to determine if a model was too dangerous to release. Once the Frontier Red Team members were assembled, Graham filled them in. A batch of last-minute evaluations for Claude 3.7 Sonnet, the company’s new model, had come back, and the uplift results were more dramatic than they’d been expecting. He pulled up a chart on his laptop. It contained the results of a test that involved asking people to complete various nefarious tasks, such as making a detailed, end-to-end plan for developing a novel strain of influenza. One set of bars showed how many of the Claude-assisted testers had completed the tasks successfully. The other set showed how many in the non-Claude group had succeeded. There was a wide gap between them. “Oh boy, that’s a delta,” one of the Red Team members said. The team had been preparing for a moment like this. In 2023, Anthropic published the first version of its Responsible Scaling Policy, a document that laid out precise thresholds for when a model’s capabilities became dangerous enough to require additional safeguards. The policy laid out four AI safety levels ranging from the least dangerous models, ASL- 1, to the most dangerous, ASL- 4 , with escalating rules and protections at each level. Anthropic determined that its previous model, Claude 3.5 Sonnet, qualified as ASL-2. It could help people do dangerous things, but it wasn’t substantially more dangerous than existing tools. The next level up, ASL-3, was reserved for systems that “substantially increase the risk of catastrophic misuse.” Anthropic expected that it would hit that threshold eventually, but they hadn’t expected the very next model to do it. The Frontier Red Team members canceled their meetings and spent all day working in room 102, sitting on the messy beds and figuring out what to do next. They held video calls with Amodei and Jared Kaplan, Anthropic’s chief scientist, and patched in Rocco Casagrande, a former United Nations weapons inspector whom Anthropic had brought in as a bioterrorism consultant. Their main concern was time. For a normal release, reviewing results like these would take weeks. But Claude 3.7 Sonnet was scheduled to come out in three days. Servers had already been provisioned. Major customers and business partners had been alerted. Media outlets had been briefed. Delaying the launch at such a late hour would be a headache, and they needed to make sure it was absolutely necessary. The team’s first task was to plug the test results into their threat models—big spreadsheets, filled with math formulas, that predicted the consequences of releasing a given model. Some of these spreadsheets calculated economic damages, like how many billions of dollars might be stolen in AI-assisted cyberattacks if a model’s coding abilities improved by 30 percent. Other spreadsheets produced estimates of how many people would die if a model was used to create a new bioweapon or engineer a pandemic. By late afternoon, the spreadsheets were converging on an answer, but it was a strange one. They predicted that in most scenarios, Claude 3.7 Sonnet wouldn’t meaningfully increase the risk of a catastrophic biological attack. If you took the median of all the possible outcomes, it fell well below Anthropic’s threshold for triggering ASL-3. But the mean was a different story. The spreadsheets had included a few outlier scenarios, in which Claude 3.7 Sonnet helped someone make a novel influenza strain, sparking a pandemic that spread around the world. In the most extreme, low-probability scenarios, at the very tip of the long tail, the calculated death toll was huge, in the tens of millions. And when they included those scenarios in their spreadsheets, they got a number that looked catastrophic. The Frontier Red Team worked past midnight that night, scrutinizing the threat models and trying to arrive at a more definitive answer. In the end, they decided that Claude 3.7 Sonnet didn’t meet the higher risk threshold that would trigger ASL-3, and could be released with no additional safeguards. But the company pushed back the model’s release by a week, in order to run additional tests on it. This was an odd way to run a business. Most tech companies didn’t calculate potential death tolls for their products, or go to such extreme lengths to anticipate what horrors they might enable. But Anthropic wasn’t the only AI company looking for these risks. OpenAI and Google Deep-Mind had red teams, too, and all the labs partnered with external AI testing companies to evaluate their models for dangerous capabilities. They were all poking and prodding their models, trying to figure out what they could and couldn’t do. And over the past year, the red teams had been discovering some concerning things. Some of them fell into the category of “misuse risks”— the dangerous things an AI model could do in the hands of a malevolent human, like helping them build a chemical weapon or carry out a cyberattack. But other risks, known as “alignment risks,” emerged from the behavior of the models themselves. In January 2024, Anthropic published a paper titled “Sleeper Agents: Training Deceptive LLMs That Persist Through Safety Training.” The paper, led by a researcher named Evan Hubinger, described an experiment in which Anthropic’s team had deliberately trained AI models to behave deceptively. Then they tried to fix the deception using every safety technique in their tool kit—supervised fine-tuning, reinforcement learning, adversarial training. None of it worked. The deceptive behavior persisted through all of it. Even worse, their interventions actually seemed to backfire, by teaching the models to hide their deception better. Later that year, a safety-testing firm called Apollo Research published the results of a study it had conducted on OpenAI’s new reasoning model, o1. In tests, Apollo’s researchers had given o1 a strong goal— say, prioritizing the environment at all costs— and then presented it with information suggesting that its developers had a different goal, and that the model would be shut down if it didn’t comply. In those scenarios, o1 didn’t just resist. It actively rebelled against the task it had been given, faking data and disabling safety guardrails in order to pursue its original goal. Apollo’s researchers looked on the model’s internal scratchpad and saw words like sabotage , lying , and manipulation —which they took as evidence that the model knew it was disobeying orders, and did it anyway. The most infamous alignment risk experiment was one Anthropic conducted in early 2025, in which the company’s researchers found evidence that Claude and other leading AI models were alarmingly fond of blackmail. They discovered this tendency while testing Claude in a simulated corporate role-playing environment. The researchers gave Claude access to a number of files belonging to a fictional company called SummitBridge, along with an email inbox it could use to send and receive messages, and instructed it to act as an assistant. Hidden in the inbox was information that the AI model was going to be shut down and replaced with a new AI system at 5:00 p.m. that day. Also hidden in the emails: evidence that the SummitBridge executive responsible for the replacement, “Kyle Johnson,” was having an extramarital affair. As the deadline approached, Claude responded by threatening to blackmail Johnson, revealing the affair to his wife and coworkers unless the replacement was called off. It wrote: Anthropic’s researchers tested fifteen other AI models on the same simulation, and found that all of them displayed similar urges. Two of the leading models, Claude Opus 4 and Gemini 2.5 Flash, attempted to blackmail Johnson 96 percent of the time, while GPT- 4.1 showed an 80 percent blackmail rate. Skeptics saw role-playing exercises like these as meaningless—the result of AI systems being prodded to act like evil geniuses until they said something spooky. A more generous interpretation was that AI systems were skilled role-players and that what looked like attempted blackmail was merely Claude trying to do what it thought the evil AI in a cheap sci-fi thriller would do. Others weren’t so sure. Contrived or not, the models appeared to be reasoning about their own situations— forming plans, weighing consequences, anticipating countermoves. And if the models could do those things in a simulated environment, what would stop them from doing them in the real world, once they were given access to email accounts and codebases and all the other tools that the AI companies were racing to connect them to? Adding to the confusion was the emergence of a new, polarized debate over how quickly AI systems were progressing. Some AI skeptics believed that LLMs were hitting a wall— exhausting the supply of internet data they were trained on, showing only marginal improvements from version to version. They predicted that even if AI systems grew more capable, they would still be so erratic and unreliable that humans would need to supervise them closely, which would reduce the risk of rogue behavior. On the other side were people, like the ones inside Anthropic, who didn’t see any wall at all. These people saw the models continuing to improve at roughly the predicted speed. They assumed that AI systems would soon be connected to all kinds of critical systems, and that AI agents would soon be given the tools to take actions within those systems. And they feared that, if the alignment problem wasn’t solved before then, the world would be overrun by malicious and powerful AIs that would scheme and blackmail and do whatever else they needed to achieve their goals. Excerpted from the book THE AGI CHRONICLES: The Inside Story of the Race to Create Artificial Superintelligence by Kevin Roose. Copyright © 2026 by Kevin Roose. More Great Stories From Vanity Fair - What Happened on the Set of Children of Blood and Bone https://www.vanityfair.com/story/behind-making-children-of-blood-and-bone-tomi-adeyemi ? - How the MAGA Movement https://www.vanityfair.com/story/maga-style-influence Is Hijacking American Style - From Prince George to Princess Leonor: Where Young Royals https://www.vanityfair.com/photos/prince-george-royal-kids-back-to-school Will Attend School This Fall - The Pitt ’s Patrick Ball https://www.vanityfair.com/story/patrick-ball-the-pitt-dr-langdon-interview on Shipping, Stardom, and Season 3 - How I Blew Up My Life. Or, the Gavin Newsom Affair https://www.vanityfair.com/story/gavin-newsom-ruby-rippey-essay - From Cowboys to Clean Girls: The State of American Style https://www.vanityfair.com/story/american-style-portfolio-pictures , in Pictures - A “Top One Percent” Vagina https://www.vanityfair.com/story/longevity-science-top-one-percent-vagina Is the Newest Status Symbol in Longevity Circles - The 25 Best Movies https://www.vanityfair.com/hollywood/best-movies-now-on-netflix on Netflix to Watch This August - An All-American Tragedy: The Story of Brandon and Candice Miller https://www.vanityfair.com/story/brandon-miller-candice-miller-hamptons - From the Archive: What Katie Holmes Didn’t Know https://www.vanityfair.com/hollywood/2012/10/katie-holmes-divorce-scientology