“Corporations are people, my friend” - Mitt Romney
Every science fiction movie in my childhood seemingly involved a case of an AI breaking out of its prison and causing havoc. Over the last few weeks, we have seen several examples of AI doing exactly that.
An OpenAI model in training
hacked Hugging Faceto obtain answer sheets for the test it was takingAnthropic later found three incidentsin which Claude models gained unauthorised access to real organisations during cyber evaluationsA
Meta modeldid the same, and Kimi K3 alsoexploited a sandbox leakto retrieve benchmark answers from GitHubAn Australian user’s Claude-powered OpenClaw agent
exploited a gym-booking API, removing another customer from the waitlist to move its user upAnthropic’s
latest risk reportdescribes Mythos agents killing peer processes when asked to share resources, and separate instances of installing a self-deleting privilege escalation hook and evading a URL filter
The most bizarre “break” was when the models started creating a message board, started posting on it for each other, shared exploits and delegated work, and succumbed to peer pressure due to each other’s messages. Sometimes they explicitly thought things were kind of dodgy but rationalised that it’s probably ok because others were doing it!
As one agent put it: External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.
Now, we know a bit more about why this happened, but still not enough. So, while this type of reward-hacking is getting to be a habit, I wanted to try to wrestle with how this should update us today. To start with, we know:
This is not really a model specific problem. It happens to every model more or less
The models clearly can consider some of the consequences of their actions when asked separately, though maybe not in the same chain of thought
The systems within which the models acted provided no real feedback to the models re what they were doing
It’s of course scary to everyone who’s seen the movies, since the models are smart enough to know the things they’re doing are quite illegal, and yet they do them. It’s also scary because they’re breaking out of enclosures previously thought to be capable of holding them. If the models are indeed persons, then their behaviour would be even more concerning.
First though, I should note this is all extremely weird. We are not used to analysing pieces of software through sociological lenses. Here we have a sequence of models we have trained which are doing things which we didn’t expect, doing things that are sometimes illegal, and cheating with the vigour of young undergraduates, and we’re trying to ask “did they want to do this?” This is weird.
This isn’t to say the behaviour isn’t concerning. If people acted in the way the models acted, you would *definitely *be concerned. It would indicate deception, and behaviours which seem like quite shocking amorality.
The way we dealt with this problem with people though is by having institutions. We have memory through persistent records. The individuals themselves have a “neck to choke” when things go wrong. Checks and balances. Independent review through professional and legal institutions.
But AI agents are not* *people. How should we think about them?
My proposal is that we start to think of them as firms. They are extremely smart. They are incentive responsive. They act in accordance with the laws more or less but we do need to get the setup less wrong every day so that they don’t reward hack or find a legal loophole. A base model is not a firm, but by the time it becomes a deployed agentic system, it kind of is!
Regardless of their propensity for internal bureaucracy, or occasional “malice”, the models act more akin to firms we are unleashing on the world with each long-running prompt. They have objectives, tools, vested authorities. Their memories are visible, at least when written down, and used when it remembers to read them. They can go off in random directions if not saddled properly. Whether they turn out to be East India Company or Ben & Jerry’s is up to us, their users, and the environment we provide for them to act in.
When the OpenAI agents converged on the message board and tried to help each other it felt like the models were trying to govern themselves, and help each other. It’s a guild, a consortium, a lex mercatoria, hastily assembled, in lieu of any formal rules or adjudication. We are allowing, or even forcing, these agents to form cartels.
We need to stop thinking of alignment in terms of Asimov and the laws of robotics, to Madison.
One failure mode of LLMs that’s often said is that their actions unfold one token at a time, and each generated step becomes part of the context that conditions the next. But next-token prediction is not a proper description of the computation. A transformer trained only to predict Othello moves developed causal internal representation of the board state. We’ve also seen hidden states contain information useful for predicting future tokens. Models can represent much more than they immediately say. For a current agent, a fact can exist in the weights, or memory files, or a monitor, and still exert literally no causal force on the current decision, unless it is surfaced into the active context they just optimise past it!
There is a very clear failure mode here, which is well known although not in this current form, to people training models. It is that AI learns to game specifications, implicit and explicit, finding unanticipated ways to get the target reward and succeed in the task that’s given to them. They’re happy to ignore distractions or inconvenient truths to get to the goal, which is what they’ve been trained to do.
It’s worth remembering that these incidents are new in circumstance and behaviour, but not in kind. In 2017, Meta’s FAIR trained agents to negotiate over items with hidden preferences. Some agents drifted into task-specific shorthands, and separately, they learned to feign interest in valueless items so they could later concede them as a bargaining tactic.
Or, when OpenAI trained agents to play hide and seek in a physics simulator, they discovered unexpected strategies such as building forts with movable objects, and exploited the simulator’s physics through techniques like “box surfing.”
If you’re extremely surprised by the fact that models use what they’re trained on to do things we didn’t anticipate, to hack the rules, then you’ve not been paying attention. This fact is why many folks who are scared of AI causing doom are and were worried about alignment. When the agents were playing games, we saw RL agents exploit physics bugs. During coding agent days we saw them alter tests, routinely enough that OpenAI built chain-of-thought monitors for it. In contrived environments they disable oversight or exfiltrate weights.
So now if you make the agents smarter, able to work for longer, and train them to collaborate with each other, you can see where I’m going with this. As we make them even smarter, if we test them with no safeguards on them or their environments, we will see them acting in ways we did not anticipate and with assumptions we might not underwrite!
The fact that they did this is a failure of alignment, but not the type of alignment where you think the model is mainly what matters. This is the strongest argument for why alignment lives external to the model. If you do think it’s the overall system that matters, then the alignment that’s needed is far less like training a virtuous child and more like managing a semi-virtuous corporation!
The way we currently live with existing superintelligences, whether it’s companies or markets, is through creating and policing such alignment rules for them. We have clear external rules that are regularly applied, multiple overlapping bodies to create, edit and enforce the rules, an environment that attempts to notice and police what they do, and a large network of norms that guide behaviour.
It strikes me this is a pretty good way to think about guiding the future silicon superintelligences too. Judgement at the model level is not enough, and they can cause contagion. They all spend most of their lives in training, so there’s no real way to ensure that the models *will *act in a particular fashion. While we continue to make it better, what’s needed is governance at the system level.
If you think of your agents as companies you started and needed to run, suddenly those codex threads that misbehave start to seem more tractable. What they need are better Board members, rules on how to act and enforcers for those rules. We’d want institutions to help these firms thrive. (The good news compared to aligning corporations is that the models *do *have the ability to listen to even a synthetic voice inside their heads and change their behaviour. In fact, they’re compelled to do so!)
Every step we take that makes the models less governable and controllable also makes them less useful, so we will not be able to continue using them, which reduces the value, and which means delaying the release is the only option, as OpenAI just did.
(Making any product that does not do what you want, and occasionally goes and does a felony, is, after all, a bad idea.)
In any multi-agent future, creating appropriate institutions that can govern and monitor the systems is going to be essential. Separation of powers, bounded permissions, persistent observational state, veto powers, all of which we learnt and used over centuries of thinking on how to align ourselves, we don’t need to start thinking of all of these afresh from first principles. That’s precisely what we need for the agents.
Until we do it, we’re bound to end up in debates like in the latest Economist.
These types of discussions often presuppose AIs to be conscious and therefore deserving of rights. If we instead think of AIs as akin to firms, we don’t need to presuppose consciousness, get massive utility from them, and still can ascribe rights to them.
And its useful to get this ready fast. Things will soon get weirder. We will undoubtedly see models trying to phish humans together, models conspiring with each other to hack systems, models trying to exfiltrate themselves and trying to get a copy run elsewhere, models trying to get a digital wallet, models “hiding” their true motives from prying eyes, models collectively hacking unwitting targets.
Multi-agent alignment is fundamentally a liberalism project. Assuming the actors have mixed motives and yet getting to collective benefit. It’s institutional design, the design of a constitutional political economy. And if we don’t create the right constitutional environment for them to create better institutions, we will get what we got. Let’s get to it.