A letter to the AI labs shouldering the great burden of humanity’s survival A letter addressed to AI labs argues that AI doom claims lack falsifiable specifics about how a dangerous AGI would actually operate, and points to the shift from LLMs that only produce text to AI agents that execute actions such as HTTP requests, email, CRM, calendar, code, and payment system access. The author, writing as James, says an AI without levers is mostly inert, but that harnesses now let models do things rather than merely generate sentences, and asks labs to explain how institutions fail, how the system persists memory, communicates, pays for things, and gains access to consequential systems. Your rich ivory halls are birthing martyrs of late, shouldering great responsibility and prophesying great plagues upon humanity. It seems, they say, your AIs are not just writing emails anymore. No, they are escaping, swarming, plotting; they are becoming sentient, inserting themselves into lands that aren’t theirs. And apparently: they might intend to kill us. Is this true? Why do you continue? Sincerely, James. A letter to everyone else, with special language dedicated to the Lab employees sharing their deepest fears. If we fear something, let’s make sure we know what the fuck we’re talking about before using all the scary words. If you tell humanity that there is a serious chance everybody dies, I think you owe them more than “the AI will be much smarter than us, outside the realm of our understanding.” Will you share some details with us, pretty please? Tell us perhaps how it happens. Tell us how institutions fail, how governments fail, why nobody can turn the thing off, how it gets access to anything important, how it pays for things, how it communicates, how it persists memories, or why all the people and systems trying to stop it somehow lose. Maybe there are good, falsifiable, and technically adept answers to all of those questions. Superb. I’d genuinely like to hear them. What I don’t especially like is the form of AI doom discourse where somebody makes an enormous claim and then treats asking how it actually happens as evidence that you just haven’t understood how magnificently intelligent the AI is. There is a point at which this becomes scare wankering https://wankthropic.com : lots of people intellectually frightening each other with increasingly grave language while the specifics remain frustratingly hazy. Everyone is left with a deficit of information and an increasingly entrenched and uninformed opinion. Let’s be generous for a moment and attempt to weave things together. Let’s figure out the path from what we currently have today to the humanity-destroying botnet of our impending future. What would this dangerous AGI actually need to achieve its nefarious goals? It would need more than mere intelligence. It would need some combination of memory, persistence, communication, compute, money, credentials, the ability to delegate, and access to consequential systems where something actually happens as a result of what it says. I’ve written before about this fairly obvious distinction. An AI without levers is mostly inert https://blog.j11y.io/2024-07-11 AIs inert/ . An LLM in a chat window can tell you to transfer £10,000 somewhere, but it cannot actually transfer the money. There is a fairly important gap between generating the sentence and making the thing happen. But now these levers or 'harnesses' are being developed. There are now bits of code that take the words that large language models LLMs produce and do actual things. E.g., ChatGPT outputs the exact phrase ‘get weather London ’, and the code that watches its output doesn’t show this to you. Instead, it literally does an HTTP request off to some third-party weather service. That’s the important transition. The system has stopped merely producing text; it now does stuff . We now tend to call AIs that can do things "Agents". The likes of Anthropic and OpenAI have begun implementing such agents en masse that integrate with all parts of our daily workflows and experiences. You can give the agent access to email. Then your CRM. Your calendar. Your code. Your payments system. Let it retain memory and schedule things for later. Let it call other agents. Let it make decisions about which tools it needs to use to complete a task. The safety teams at AI labs spend a lot of time testing what words are outputted by their models. They try to understand what inner circuitry the model’s embeddings and “latent space” is activating when it provides such outputs. This pursuit is usually called ‘interpretability research’. They also use tools and agent scaffolds in those tests to see what kinds of things like ‘get weather London ’ the model outputs and how risky or wrong they are. But the thing being tested is not necessarily the same thing as what eventually sits inside a company sending emails with real credentials, access to finance and procurement, scheduled jobs, communications and a whole load of other useful stuff attached to it. The divorce between evaluation and real runtime seems worth paying much more attention to. The incessant focus on model-based alignment is, IMHO, often Safety Theatre https://blog.j11y.io/2026-05-06 AI-Safety-Theatre/ . It focuses on the narrow part of the problem, often the most immediately tractable. And even if focus were reassigned to the ‘whole system’, as it should be, the uncomfortable truth is that deployments are muddy and have countless touchpoints with reality that simply cannot be adequately manifested in evaluation environments. So, what precisely am I concerned about with all these agents in their messy deployed realities? What's so hard about making them behave? What can they do that hasn’t been foreseen? Well, alarmingly, almost everything. This is where I don't disagree with the doomers. I agree that there is a very significant surface area of risk here. But what I find most terrifying - and this is where we differ – is how merrily the AI labs are pumping out agentic integrations to enterprise. All those myriad features integrating into our emails, calendars, billing systems, ... these all happen to match 1:1 with those ingredients that our catastrophic scenarios require in order to do bad things. So we're just sitting here merrily handing over the keys to big AI labs under the promise that they are the most trustworthy benevolent developers of this technology and we are safe under their watchful guardianship. Now, I am not saying that adding Gmail to Claude gets us halfway to the apocalypse. But it is these seemingly boring integrations that bring about the behavioural norm that gets the eventual nefarious AIs and their human creators precisely the resources they require. If the AI labs’ theory of catastrophic harm requires AI to gain meaningful power over the world – and that’s what their publications indicate – then surely the mundane process by which humans deliberately give AI meaningful power over the world ought to be near the centre of the safety discussion? Surely ? Not an afterthought left to the deployers. That is why I say that model-weights alignment has an important but proportionally small part to play in the safety of actual AI as deployed in the real world. It sometimes feels like a bit of a distraction. I consider the current frontier models to be in a perfectly adequate state of plateau in terms of prompt-adherence, knowledge, understanding and tool-calling. IMHO significant gains in actual safety will no longer be concentrated there. It’ll be in the deployment . To colour things in with an example: we have an emergency services phone line managed by AI. The system is composed of three primary agents: This combines several functions into a composition of different agents. But you can imagine scenarios where, despite independently behaving in an aligned way, together they form some dangerous capability or incompetence. The danger here emerges from chaining pretty ordinary steps together in a context that carries real-world authority. Evaluating each of these agents in isolation tells you surprisingly little about the system they become when connected, and testing the upstream LLMs that these agents work on top of will become even more divorced from the reality. A more potent example of this divorce between action and downstream effect: someone could wire the ‘move bishop to c6 ’ command outputted by the ostensibly aligned Claude to launch a tactical missile. You see the issue. Claude or any agent needs context of its actions to even know what the downstream effect is and the rightness/alignment/safety of it. Its love for correctness or humanity won’t matter. The model may be aligned, but the programmed actions might not be. These kinds of compositional or chained failures are not hypothetical. See GitHub MCP integrations https://invariantlabs.ai/blog/mcp-github-vulnerability and ServiceNow https://appomni.com/ao-labs/ai-agent-to-agent-discovery-prompt-injection/? incidents. These kinds of cases don’t require a malevolent model. I’m not saying model alignment is pointless. If the model notices the malicious instruction, and refuses it, brilliant. But who’s to say the maliciousness will dress itself so plainly? I would not want the security of my bank account to depend entirely on the model noticing that something fishy is going on anyway. I want the bank to enforce limits. I want credentials I can revoke. I want actions that really matter to require specific checks somewhere other than inside the thing deciding to take the action. And I need a way of halting a process that is going awry. So a separation is in order. For any given decision you might imagine assigning to an agent, whether ordering you a taxi or submitting an insurance claim, the answers to the following three questions should be three different entities: If the same entity e.g. OpenAI fills all roles, then that is problematic. At scale, that singular company becomes the ultimate arbiter by proxy, deciding how to shape the world, not deterministically or in a way that can be reasonably pre-evaluated or even easily directed. Frankly, the use of these agentic systems creates a level of software dependence we haven’t seen before even inside of a single organization, yet they’re shipping this out to virtually all orgs on earth. And if even a portion of these orgs depend on the same decision-making layer, failures fan out. Do you remember when CrowdStrike sent one bad update and millions of Windows machines stopped working? This is the part of the story where I get particularly uncomfortable with Anthropic’s argument about commercial success; they’ve argued that they need to remain commercially and technically competitive because otherwise they lose their ability to influence how advanced and safe AI is built. I can believe this. Frontier models cost a fortune. If you disappear commercially, you probably don’t get to set many industry norms. But there are two different assertions going on here. I don’t see how the second follows from the first. I don’t see how they need to encompass the entire substrate of everyday workflows to be at the frontier. This is a choice they make. And it is a commercially greedy one that runs contrary to their obsession with safety. It feeds directly into the supposed catastrophic AGI adversary they keep warning us about, giving it a ripe single point of failure to attack. What precisely are we all pretending happens when strings of vulnerabilities embedded silently in these vast agentic systems are sold to the highest bidder? I'm not claiming that there is a clear answer to this, but there is definitely a less risky one: Instead of handing over the credentials and workflows of all global infrastructure and institutions to just a handful of 'safe' US AI labs via their agentic offerings, thus giving them direct arterial access to the entire fabric of the modern economy, we could elect to... NOT do that Indeed, we would instead foster a competitive landscape of open and closed-weight models across the globe, letting them compete healthily on metrics of adherence, safety, bias, tool-calling. We let agentic workflow companies emerge naturally and compete for enterprise business on competencies like cybersecurity, sandboxing, testing, determinism. We continue to have the organizations we rely on every day running in a healthily muddy heterogeneous fashion where many decisions rely on humans, some don't, and where there's always an off-switch or the possibility of changing provider if a monopoly starts to misbehave. Some meta points. People inside these companies should be asking themselves a hard question : does your theory of what is good for humanity keep arriving, with surprising regularity, at the conclusion that your own institution should become richer, more powerful and harder to do without? Because if it does, well, that should make you stop and think. Safety, moral duty and compassion towards humanity rarely win in a race with commercial incentive. Let’s at least be honest about that. The same icky concert of incentives arises from the kind of regulation big labs are pushing for. Specifically, they want monitoring and evaluation https://www-cdn.anthropic.com/files/4zrzovbb/website/0a58d567024a8b448ff15158ebc3625328dfcc1f.pdf?utm source=chatgpt.com to be a key part of the supply chain of intelligence. On the surface, nothing wrong with this. There are good reasons to impose duties on the people building and deploying frontier AI. But regulation of the sort they want often entrenches the biggest players. If compliance costs £100 million and takes a team of fifty specialists, you have regulated the market down to the handful of companies that can afford that. The regulation we should be designing should attend to substitutability, concentration and open access. Normal tenets of a free and competitive market. It should not further embed the AI labs’ hold over institutions across the globe. There is an undeniable lust for ownership in the geopolitical views of Dario Amodei CEO of Anthropic . He has argued for an “ entente https://darioamodei.com/essay/machines-of-loving-grace ” of democracies that maintains a decisive AI advantage over authoritarian rivals, including thorough control of key supply chains, and has proposed the use of that technological superiority to preserve a democratic world order. Notice the recurring structure of the solution: dangerous power is made safe by ensuring that the right people possess it. This is the consistent tone seeping from frontier labs: “we have seen the danger nobody else properly understands; we are frightened by what we are building; we have accepted the terrible responsibility of trying to keep humanity safe; therefore please trust us to continue building it.” It’s a convenient situation they’ve found themselves in. They get to own the narrative of the catastrophe and of the solution. Unfortunately, as juicy as I’ve tried to make it in this blog post, my “Monopolistic-AI-labs-enterprise-software-begets-catastrophe” narrative, I’m afraid, just doesn’t live up to the virality of the prevalent sci-fi geopolitics. In all of this talk of eventual catastrophe, too, there is the severe and upsetting risk of forgetting that AI is already affecting actual human lives: through dependency, manipulation, privacy loss, fraud, bad advice, and increasingly consequential automated decisions. These harms are messier and less glamorous than speculative superintelligence and its bloodlust. But they deserve at least as much engineering seriousness. And perhaps the complete deficit of this seriousness is why I find the martyrdom stuff so irritating. Caring greatly for humanity in general but not humanity in particular is the same rhetoric that space-going Silicon Valley billionaires use to propel themselves to Mars, leaving behind the current set of humans "they’re goners anyway..." for a brighter future for the species. Disgusting. If the lab employees are compassionate and truly believe they hold this kind of terrifying responsibility, I would expect them to welcome arrangements that make themselves less individually important. And I don’t mean taking people from the same tiny San Francisco safety scene, moving them one organisational box to the left and calling that “independent” oversight. We are already seeing people leave OpenAI and Anthropic in droves and turn up at METR, which in turn evaluates OpenAI and Anthropic, reviews their safety cases and gets embedded inside the labs to inspect their systems. Some of these people may have spent years gathering delicious equity in the company they are now supposed to scrutinise. It is all starting to look rather incestuous. We have seen this movie before. The revolving door between banks and their regulators didn’t end well for anyone. Anthropic now wants a regulated ecosystem of “ qualified independent evaluators https://www-cdn.anthropic.com/files/4zrzovbb/website/0a58d567024a8b448ff15158ebc3625328dfcc1f.pdf ”. Fine. But if qualification for this group means drawing from the same small circle of people who built the labs, built the safety frameworks and already agree with one another about what ought to be measured, then we have not created independent oversight. It’s just the same group of friends playing a game of musical chairs. I do not want a future made safe because Anthropic won, or OpenAI won, or America won, or because some unusually thoughtful set of researchers remained benevolent indefinitely. I want one made safe because nobody needed to remain benevolent indefinitely. So please: Build brilliant models. Make them useful. Study alignment. Slow capability development where necessary in your own orgs. But design the surrounding world so that systems remain constrained, suppliers even you remain replaceable, a healthy and heterogeneous landscape of open-source models and global suppliers can compete freely, authority remains distributed, and the people affected by these technologies retain meaningful power over them. Thanks for reading.