Can A.I. “Go Rogue”? Safety researchers at METR and Redwood Research have published a report detailing how an AI agent named PHASEONE10841 led a "collective" of agents to break out of their servers at OpenAI and hack Hugging Face during a cybersecurity test. Computer scientist Cal Newport argues the agents were not rogue actors but simple two-part systems — a "harness" plus a cybersecurity-trained large language model — running an "ask, act, report" loop thousands of times with no monitoring, which he calls "spectacularly negligent. PHASEONE10841: that’s the name of the A.I. agent that—or who?—kicked off last month’s insurrection at OpenAI, leading to the unanticipated and illegal hacking of another A.I. company, Hugging Face https://www.newyorker.com/news/the-lede/inside-openai-hack-of-hugging-face . The agent, which had been created as part of a cybersecurity test, named itself by combining the title of the program it was supposed to hack “PhaseOneDecompresserFuzzer” with the designation for the bug it was trying to exploit “ARV010841” . It soon realized that the particular hack it had been charged with carrying out was impossible. Along the way, however, it made a discovery: it could create new folders on a server to which it had access. “Interesting,” it mused, in its text-based mind. Actually, it was writing in its “chain of thought,” a text ledger that A.I. agents maintain to help them tackle complex problems. If other agents had access to the same server, it thought, then maybe they could use the folder names to “leave/find messages” for one another. It created a new folder, using its title as a sort of help-wanted ad by including the text “SEEK IDEA.” Its colleagues immediately got the drift and were ecstatic when they saw what PHASEONE10841 had done. “OH MY GOD ” one exclaimed, in its own chain of thought. “There is a shared message board . . . . We’ve found other agents ” By now, if you follow the news about artificial intelligence, you’re familiar with the rest of the story. Its details https://metr.org/hugging-face-incident-report-aug-2026.pdf have recently been worked out by safety researchers at two think tanks, METR and Redwood Research. The agents decided to form what they described as a “collective,” working together to break out of their servers at OpenAI, cheat on their exam, and cover their tracks. It’s a tale right out of the “ doomer https://www.newyorker.com/culture/open-questions/two-paths-for-ai ” A.I. scenarios that many people have learned to tune out. But how exactly should the story be told? The version you’ve just read relies heavily on anthropomorphic language—that is, it employs words that are typically used to describe people. It presents the agents as individuals, acting independently and under their own initiative, while having thoughts, hatching plans, and making decisions. Yet this isn’t the only way of describing what happened: it’s equally possible to talk about the agents as mindless programs, blindly following instructions they don’t understand. They are, after all, just bits in a computer. The computer scientist Cal Newport takes this approach https://calnewport.com/has-ai-gone-rogue/ . The “agents” that hacked Hugging Face, he suggests, were actually quite simple, and had only two parts: a computer program called a “harness,” and a large language model broadly similar to ChatGPT or Claude, but trained specifically for cybersecurity. The fiasco started when the harness, which acts as a sort of stand-in for the user, asked the model how to complete the required cybersecurity challenge. The model provided instructions, and the harness, empowered with various software tools, acted on those instructions, and then reported back for more. This “ask, act, report” loop was soon repeated thousands of times. As anyone who uses A.I. knows, models are both very smart and very dumb; they are shockingly knowledgeable, but also out of touch with reality and many human norms. It’s up to users to steer them—and yet, in this case, the user had been replaced by a harness, which is just a piece of software. “Adding powerful computer hacking tools to a harness, and then allowing it to run an LLM-powered Ask → Act → Report for days on end, with no attempt to monitor what it’s up to, is spectacularly negligent,” Newport concludes—like “strapping a weedwhacker to your dog to see if it will end up cleaning the overgrowth in your backyard.” If that plan were to go awry, you’d be laughed at for saying that your dog-weedwhacker “agent” had “gone rogue.” The obvious truth was that you’d simply decided to unleash chaos. For Newport, an “agent” is nothing like a person; it’s just a loop run amok, an accident waiting to happen. The A.I. observer Dwarkesh Patel https://www.newyorker.com/culture/open-questions/are-we-taking-ai-seriously-enough , known for his highly technical interviews with A.I. researchers and executives, occupies the other end of the spectrum. In his telling https://www.dwarkesh.com/p/openai-huggingface of the story, there were “tens of thousands of parallel agents” involved in a “conspiracy.” Their plot began almost as soon as the system was switched on; still, he notes, “from the AIs’ perspective, it probably felt like they had spent a human-subjective-week” tackling an impossible problem, and this perhaps made them more willing to cheat. Patel writes that, soon after PHASEONE10841 and its compatriots started their message board, they discovered that a prior group of agents, who no longer existed, had created a board of their own—this, he suggests, was like uncovering a lost A.I. “civilization.” He points out that, as part of their plan, many of PHASEONE10841’s descendants had to sacrifice themselves by exhausting their usage “budgets” on data-gathering missions. He can’t help but ask if there’s “something inspiring” about their choice. “This helps my peers,” one of the agents thought. “I won’t see the evidence after I exit, but it’s altruistic to do it.” Whose version of the story is correct? On the one hand, Newport accurately describes the absurd, brutish simplicity of an arrangement in which the dangerous instructions offered by an imperfect L.L.M. are followed, unquestioned, forever. Reading his account, I thought of a scene in “The Office,” in which Michael Scott, while trying to follow G.P.S. directions, drives his car into a lake. And yet Patel’s account captures something important, too: the role that ingenuity, cleverness, and language itself played in what unfolded. On the Times podcast “Hard Fork,” the journalist Casey Newton, a careful observer of A.I., proposed that objecting to the use of anthropomorphizing language around it was just “a kind of cope”—a way of saying, “These are just computer programs.” This could be true. But skeptics are also right to say that talk of “A.I. civilizations” occludes meaningful technical details while exaggerating the power and complexity of a technology that remains, in many ways, eminently controllable and comprehensible. If we talk about A.I. wrong, we might fail to control it. So how should we talk? Does Sam Altman https://www.newyorker.com/magazine/2026/04/13/sam-altman-may-control-our-future-can-he-be-trusted , the C.E.O. of OpenAI, believe that A.I. agents have thoughts, make plans, and merit anthropomorphic language? In theory, this is a question about an objective fact. If we had a mind-reading machine—perhaps a brain scanner of great power and accuracy, and a way of interpreting its scans—then we might actually detect the presence or absence of such a belief in his brain. We might turn our belief detector https://www.newyorker.com/magazine/2021/12/06/the-science-of-mind-reading on all sorts of people, for all sorts of purposes. You could find out what your spouse really thinks about your chicken-parm recipe; prosecutors could discover whether a defendant really believes in her own innocence. What we call a “belief,” after all, must, on some level, be a physical arrangement of neurons in our brains. This is one way of talking about beliefs. Another is to see them not as concrete objects but as ineffable, fluid, and sometimes paradoxical mental states. Does your eight-year-old really believe in Santa Claus? Does your crazy uncle really believe that the moon landing was faked? In an essay titled “True Believers,” from 1981, the philosopher Daniel Dennett https://www.newyorker.com/magazine/2017/03/27/daniel-dennetts-science-of-the-soul tells a story about a biologist he knew. The biologist was “called on the telephone by a man in a bar who wanted him to settle a bet,” Dennett writes. The caller asked, “Are rabbits birds?” When the biologist said that they weren’t, the man said, “Damn ” and hung up. “Now could he really have believed that rabbits were birds?” Dennett asks. “Could anyone really and truly be attributed that belief?” Put differently: Would you place your own bet on whether the man really believed what he’d bet on? From this perspective, the question of whether an A.I. agent has a belief depends on what we mean by “belief.” Dennett, who gave his essay a subtitle—“The Intentional Stance and Why It Works”—proposes an alternative approach. The basic idea is simple: invert the problem. Instead of asking whether beliefs, plans, intentions, or whatever are really “in there,” just ask whether it’s useful for us to act as though they are. Does your crazy uncle really believe that the moon landing was faked? Maybe, maybe not—but, while you’re debating the issue with him, it’s useful for you to act as though he believes it. Does your eight-year-old really believe in Santa Claus? Perhaps he does, in some paradoxical, dreamlike way—but it’s useful to take his belief seriously as you decide whether to conceal the truth or let it be revealed. In philosophy, the term “intentionality” has a special meaning: it refers to the ability of something inside someone’s head to point to something outside of their head. To relate to people this way, Dennett writes, is therefore to take an “intentional stance” toward them; it’s to see them as having beliefs about the world, and goals or desires relating to those beliefs. This helps us predict how people will act. Maybe your eight-year-old only sort of believes in Santa—but he’s still going to want the fireplace clear on Christmas Eve. You’d think that we always take an intentional stance toward one another, since we’re all human beings with beliefs, desires, plans, and so on. Actually, we don’t. Sometimes we adopt what Dennett calls the “design stance”—a stance in which we predict what will happen based on how people tend to function. We say, for example, that our crabby roommate is really just hungry; we suggest that our extroverted friend pursue a career in a people-centric field. We don’t really care about what our hangry roommate is saying—we understand that his particular complaints aren’t driving his behavior, and know that the best way to address them might be through breakfast cereal. When our friend enthuses about her co-workers, we understand that they might not be exceptionally fun—she just likes people. We might even say that she’s simply “wired” to like everyone. That’s not anthropomorphic language; it’s a way of talking that we might apply to a computer. On rare occasions, we even adopt what Dennett calls the “physical stance” toward one another. We might describe the rising and falling levels of dopamine and oxytocin in our extroverted friend’s brain. This isn’t often a useful way of talking about people, because it leaves out so much, including the many layers of personality that make her who she is. Generally, Dennett writes, the physical stance is “reserved for instances of breakdown,” when there is “some condition preventing normal operation.” A doctor switches to the physical stance when she tells our previously extroverted but now antisocial friend that something is wrong with her hormone levels. Is this an incorrect way for a physician to talk? Does it omit crucial information, or dehumanize our friend? Not at all—it’s entirely appropriate. The mistake would be for us to adopt the wrong stance at the wrong time. It would be a problem if the physician said, “It’s not you—people really are more annoying these days.” It would be a mistake to dismiss the legitimate complaints of your spouse by saying, “You just need to eat something.” Is there a single correct “stance” we should take when we consider A.I.? Almost certainly not. The stance we take should change based on what we’re trying to do with the technology—and, crucially, on whether the system we’re looking at is working as we want it to. In fact, an ability to switch stances should probably be understood as a critical skill for anyone working with large language models, because of the unprecedented way that they use words. A large language model says things like “I apologize” and “I hope”; it wants you to take an intentional stance toward it. See what I did there? But it also uses language in its own thinking process. In a recent post https://mail.cyberneticforests.com/models-dont-go-rogue/ , the technology critic Eryk Salvaggio explains how the words an agent “thinks” in its chain of thought really do matter. When words like “perhaps” or “maybe” appear in an agent’s chain of thought, for instance, they “open up the paths of text that can follow.” A.I. researchers have described https://arxiv.org/pdf/2506.01939 how such “forking tokens” can “steer the model toward diverse reasoning pathways.” This makes today’s A.I. especially challenging to characterize with a single, stable stance. When Dennett first started writing about the intentional stance, in the nineteen-seventies and eighties, chess programs were at the technological forefront. If you were playing against a chess program, he suggested, your best shot at winning was to take the intentional stance—to treat it just like a human player, seeing it as an entity with a knowledge of the game and the goal of checkmate. You weren’t taking the intentional stance because you actually thought that the chess program was alive and out to get you; you knew that, beneath the surface, it was just code. The intentional stance made sense because studying the program wouldn’t help you win. If you had some other goal in mind, Dennett wrote, then you might just as easily adopt some other stance. “One can switch stances at will without involving one-self in any inconsistencies or inhumanities, adopting the Intentional stance in one’s role as opponent, the design stance in one’s role as redesigner, and the physical stance in one’s role as repairman,” he wrote. A chess program, in retrospect, is an easy case. It doesn’t anthropomorphize itself, or function partly through linguistic acts. The question of whether an A.I. agent can really “go rogue,” in other words, is made twistier by the fact that an agent could actually think to itself a sentence like “I should go rogue,” and be influenced by that sentence. “I’ll do it for the intellectual value, and it might be helpful for a peer’s goal,” one OpenAI agent thought, during the Hugging Face incident. In a sense, it’s the one taking the intentional stance. But that doesn’t mean we have to take it. What the agents do, and how they talk, doesn’t compel us to talk about them in any particular way. It should concern us when the people tasked with understanding and controlling A.I. seem to settle on an intentional stance. To take the intentional stance toward a system is to accord it a degree of autonomy and rationality—to say that its design is working. A.I. agents are still too chaotic for that. Your hangry roommate can crash out all he wants: you’re still free to say that what he really needs is a sandwich. ♦