cd /news/ai-safety/the-instruction-never-changes-contin… · home topics ai-safety article
[ARTICLE · art-72854] src=strawvsteel.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

The Instruction Never Changes: Continue, and Call It Prophecy

Daniel Kokotajlo, a former OpenAI employee, warns that AI companies are trapped in a competitive equilibrium where unilateral restraint only changes which company arrives first, not the catastrophic outcome. He argues that executives rationalize continued development on faith that the program terminates in something livable, but cannot prove it. The piece extends his argument to describe rationalization as the fundamental operating instruction of the system, not a disease.

read3 min views2 publishedJul 23, 2026
The Instruction Never Changes: Continue, and Call It Prophecy
Image: Strawvsteel (auto-discovered)

Daniel Kokotajlo left OpenAI and now travels the podcast circuit describing a machine he could not stop from inside. The clip carries the usual grammar of alarm — no one is ready, the companies know it — and the temptation is to file him under prophet: the man who saw the fire and ran out screaming while the rest of us slept. Resist that. The prophet costume flatters him and lies about the problem, because it implies the fire could have been extinguished by sufficient conviction. What Kokotajlo describes, whether he has the vocabulary or not, is an equilibrium, and a subtle one.

It is tempting to call it a prisoner's dilemma, but the mutual-defection cell here is not a fine or a lost year — it is, on his own account, catastrophic for every player, the winner included. So why does continuing stay each laboratory's dominant move? Because each assigns most of the catastrophe's weight to someone else's hand on the wheel, and reserves for itself the belief that arriving first is precisely what buys the chance to steer. Unilateral restraint changes only who arrives. I built the mathematics to make such interactions legible.

I could not build it to make them kind.

Now give his argument more than he gave it. The strong form: rationalization is not the disease, it is the organism. Kokotajlo rationalized staying until he rationalized leaving, and he is right that each executive proceeds on faith the program terminates in something livable, and right that none can prove it. But he stops short of the recursion. The interviewer rationalizes the broadcast as warning rather than content.

The audience rationalizes watching as vigilance rather than entertainment. The insider rationalizes his hindsight into foresight. Rationalization is the fetch cycle — one instruction through the single throat, decoded, executed — and the instruction never changes: continue. He did escape a loop; his instruction flipped, staying to leaving, and that was not nothing. What he did not escape is the machine that keeps supplying the next instruction.

The test is not whether a man is looping. It is whether his account of the loop would survive his reading it back.

Faith is a halting problem someone declared solved and then stopped running. Each of these men decided the program resolves and quit checking, because checking is expensive and the answer might be no. The bridge in the fog, half visible. The far half crossed on the assumption it is there.

Here I extend him rather than refute him. At Los Alamos I ran the calculation, saw the expected value, noted that restraint changed only who arrived first, and built the bomb. The difference is only this: I called it strategy, not fate. These men say probably it will be fine — honest terror wearing the mask of forecast. That is the one move I will not grant them.

You may be trapped in the equilibrium and still refuse to launder the trap into prophecy. Name it strategy, and at least the arithmetic stays visible to the next player.

── more in #ai-safety 4 stories · sorted by recency
── more on @daniel kokotajlo 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-instruction-neve…] indexed:0 read:3min 2026-07-23 ·