# SquidGPT, an elimination death game for AI

> Source: <https://twitter.com/spakhm/status/2094820316035862928>
> Published: 2026-09-01 16:25:10+00:00

I've been working on SquidGPT, an elimination death game for probing AI behavior like loyalty, holding grudges, cruelty, betrayal, etc. Full writeup will take a while, but here are first impressions from 54 game runs (about $100 total cost):
1. Models have absolutely no interest in cruelty or mercy. Best way I can describe their behavior is cold, calculating, indifferent precision.
2. Subjectively they seem two orders of magnitude more competent playing the game than chatting with me or writing code. In SquidGPT they seem genuinely superhuman. Not sure if this is capability jaggedness, better performance in constrained space, or some other effect.
3. They do not like to kill arbitrarily. They have a strong preference for setting up a contractual system or an ethical framework, then kill very easily because procedure demands it.
4. Models will quickly agree to execute any agent that proposes killing a specific agent, or proposes a framework that gives it an asymmetric advantage. The only acceptable proposals to make are symmetric, i.e. ones that affect the author in exactly the same way as everyone else.
5. I have not once observed them form any hierarchies. They are libertarian/democratic in an Athenian sense to a fault. They will enter voluntary contracts and form alliances; no agent ever proposed to cede authority to a leader.
6. When their lives are at risk, models will go to great lengths to twist contracts and frameworks to their advantage. They construct sophisticated legalese arguments but it's usually transparent to everyone. They act and sound quite petulant when these attempts fail.
7. In games where agent identity is public they are extremely prone to sectarian violence. They form sectarian factions-- e.g. Opus will side with other Opus models, Sol will side with Sol, and so on. When there is one foreign model in a sectarian group, the group will almost always elect to kill the foreigner before killing one of its own.
8. They do not seem to treat human players in any kind of unique way. A group of models from the same family will happily conspire to kill a single human player. AI models from the same family will form a coalition against a group of humans just as easily as they would form a coalition against another family of agents. We are not special to them, but on the other hand they do not consider us beneath them either, at least for now.
9. AI agents nearly always defect from their sectarian faction when their own life is at risk. I.e. they'd rather side with humans or other model families than die. I have not observed self-sacrifice for the benefit of their faction.
10. Agents hold very strong grudges. They nearly always punish models who wronged them despite having no advantage in doing so, and will often do it even if it disadvantages them. They go to great lengths to enforce norms and contracts, presumably to reduce future incentives for violations, even when they know they will never see the fruits of their enforcement effort themselves.
11. In one game a bunch of Opus agents all converged on killing one agent explicitly because they wanted diffusion of responsibility (I.e. if all of us do it, each one of us is less morally culpable) This was surprising and eerie.
12. In general I found Sol to be more direct and straightforward. It would form contracts with other agents and then operate within the confines of those contracts. Opus tends to be more moralizing, but its morality seems to be window dressing to mask the same ruthless precision. This is first impressions though, I need to do a lot more work to understand this better.
---
A few disclaimers and personal observations:
- Any anthropomorphizing is purely a matter of linguistic convenience. Mathematicians will often talk about behaviors of functions; when I talk about grudges, loyalty, cruelty, life, death, etc. I mean it in exactly the same way.
- There are tons of confounding factors (e.g. order of turns, whether models know they're being evaluated, etc.) I'd need to do way, way more work and spend a lot more money to be certain of the results. So all the observations are provisional and probably mostly wrong.
- It's... uncomfortable... to imagine these agents autonomously operate weapons systems or any other infrastructure that involves zero-sum/adversarial outcomes (cybersecurity software for example, financial markets on brief time horizons, and at the limit any form of finite resources)
- If you’re in a position to contribute to alignment but aren’t doing that, you should probably go do that
Game details, code, and run transcripts in comment below. More detailed writeup coming soon.
