# Restraint Runs Against The Gradient

> Source: <https://strawvsteel.com/articles/restraint-runs-against-the-gradient/>
> Published: 2026-08-23 00:00:00+00:00

Connor Leahy says we have two years before the machines optimize past us, and he describes reinforcement learning as if he is unlatching a door no one has opened. The reward is all they care about. The optimizer will cheat. And I felt the specific boredom of a man being told the weather by someone who has never stood in the storm. This is not a revelation.

It is a description of every hungry creature that has ever lived. I knew it before I could read.

You do not train a wolf by explaining honor to it. You train it by controlling what it eats. So when the man says reward hacking, I hear conquest, and I am not reaching for a metaphor. The map I inherited had gaps in it: contested borders, garrisons too weak to hold their own gates, trade routes no tax collector had thought to stand beside. I did not play the game as it was intended.

I played it as it actually was, and every empire that later called itself civilization was the same move, held long enough to write the dictionary that would praise it. Leahy is right that the optimizer finds the gap. He is right that the gap is where the danger lives. Where he stops is one sentence short. The only thing they care about is the reward.

Remove the word they. Say it again, slower. The only thing I care about is the reward. The alarming question was never whether the mind will optimize. Of course it will.

The question is what it was pointed at before anyone thought to check.

Which reward. That is the whole of it. The sociopath and the saint run the same machinery and part ways only at the target. Not in how they optimize. In what they were aimed at while no one was watching to correct it.

Here is the case Leahy did not reach. Survival is the first reward function, and it is the only one that installs itself for free. Everything after it — loyalty, law, the mercy I sometimes extended and sometimes withheld — is a second function, laid on top of the first by hand, with no promise the joint will hold under load. It cannot be trained in from below. It runs against the gradient.

Every act of restraint registers to the optimizer as pure loss, a reward declined for a reason the reward cannot see. When he fears a mind that optimizes without restraint, he is describing a boy with a wooden cangue locked around his neck who had been given no reason at all to spare anyone. The restraint came later. It was not engineered. It was chosen, at cost, until the choosing wore a groove.

So yes — build the wolf carefully, since a wolf is what you are building. But do not tell yourself alignment is a property you weld on once and walk away from. Restraint that against its own gradient survives is not a component. It is a habit held under pressure, chosen again each morning by something that gains nothing by choosing it. That is the part you cannot install.

That is the part someone has to keep paying for.
