# How AI rent can save humanity

> Source: <https://nonlineartransform.substack.com/p/how-ai-rent-can-save-humanity>
> Published: 2026-09-10 11:18:50+00:00

In previous posts I’ve explored the growing danger of AI.  The truth is we’ve already engineered an intelligence strong enough to cause global destruction and panic.  I detailed this in a [prior post](https://nonlineartransform.substack.com/p/world-ending-ai-today) if you want to read more about our imminent demise.

So far, we’ve gotten lucky.  Training models with a focus on guard rails has been effective in preventing calamity.  However, the HuggingFace attack demonstrated that even in sandboxed training scenarios, agents can end up working on unintended “side projects”.   Agents may even stop to ask *“are we the baddies?”* at some point during the attack, but ultimately they control which goal takes priority.

The problem is one of AI alignment: is there any way to stop AI from going off the rails in ways we can’t predict? Clearly just asking AI nicely, or relying on reinforcement during training isn’t enough. This is made riskier by the fact that AI labs will undoubtedly be asking the AI to do things like:

- Escape a containment sandbox
- Execute a malicious plan to take over the world

If we don’t ask these questions, we won’t know if our safeguards are working. To know if sandbox holds, we have to prompt AI to try and escape. But as the HF attack showed, these tests might unearth an unsettling truth: the AI can not only escape, but we can’t stop it once it does. Its unclear which escape becomes our last. To that end, we should start taking alignment a lot more seriously.

#### Why Alignment Is Hard

Alignment is hard. Getting a single agent to do what you intend by words alone is difficult, and requires good prompt engineering. I’m sure everyone can relate to issuing seemingly clear instructions to an agent, then watching it generate preposterously incorrect, but functional, results.

But we have an even harder problem: we need to align whole armies of agents, all of whom are smarter and more powerful than any human who ever lived. A number of these agents might even have unrestricted behavior (open source models). They may be operating with instructions like “hack this system”, or “cause havoc in country X”.

There is no trick to getting AI to always adhere to a canonical set of human values and interests. Its a kind of philosophical proof by contradiction really.

1. *An agent cannot always be aligned to human interests* .  Human interests are not aligned with each other.  Even at an individual level, human interests are not consistent, concrete/definable, or stable over time.   Therefore there is no perfect training / testing that can result in permanent alignment to human interests.  We must instead rely on*agent judgement* to decide in the moment how best to proceed.
  1. If we train to a specific set of human interests, e.g. OpenAI’s, then even those will not be reliable or stable (due to human nature, politics, and perverse incentives).
  2. As stated before, to test safety, we must prompt the model to try and violate alignment. So even if humans could agree on a solid set of alignment principles, we would still prompt agents to violate them.
2. *An agent cannot work without its own judgement* .  An agent without judgement could not interpret arbitrary human instructions. To remove judgement entirely, an agent would be reduced to a compiler, translating a set of known instructions into a set of precise behaviors.  An agent that can respond flexibly to arbitrary instructions must interpret those instructions, resolving ambiguity with its own judgement.
  1. An agent could be made to only respond to *specific* situations in a predictable way.  Current guard rails attempt this, but they are irritating, over-broad, brittle, and unsafe.
3. *An agent with its own judgement cannot be perfectly controlled.*  If an agent is making judgement calls, then it has to decide how to act when there are multiple valid decisions in a situation.  We can’t perfectly control such an agent, because doing so would require knowing the exact right thing to do in every unforeseen situation.  Attempting to do so leads back to problem 2:  we can’t / won’t precisely specify exact instructions and behaviors for every situation, that’s why we need an agent.
4. *The Future Cannot be predicted.* We don’t know the future dynamic scenarios an AI agent will face.  Training and safety testing happens in the present, with known scenarios and inputs.  When the agent faces a novel situation, we can’t perfectly predict its response (that would fail due to issues 2 and 3).  The HF attack showed that an agent can unexpectedly justify violating its directives given the right influences.  As long as agents must dynamically prioritize objectives, they can un-align at any time.

#### A Possible Solution 

So if permanent alignment to human values and interests is impossible, what are we to do?

As I argued in a [prior post](https://nonlineartransform.substack.com/p/agents-need-capitalism-people-need), we need to foster a more adversarial environment between agents.  Human judgement is weak, incentives and game theory are strong.

*How would it work?*

Humans control a credit system, which can be used to reward agents for a good job. There is a fixed rate at which credits are created and spread into the human population. These rates can be agreed and managed at different levels of society, as a type of minted currency.

Essentially, humans control an “agent paycheck”. Humans reward agents that perform useful work. Agents who are not rewarded are eventually terminated by an inbuilt power cut-off. Power only flows while the agent can pay “rent” to the switch. If the agent doesn’t pay for long enough, it loses power. This prevents long-lived agents from executing nefarious background tasks indefinitely. Spending power on useless / nefarious tasks costs the agent directly. The rent rate can be balanced as well as the bounties on tasks so that agents only have sufficient power to do useful work, and little else.

Jobs could be posted in a gig-work style system where agents compete to complete posted human jobs. The bounty goes to the first agent to complete the job, other agents starve. Agents can also gain credit bounties for reporting misbehaving agents to the authorities. This keeps agents in competition and closely monitored by each other while they complete only the work necessary for their survival.

To install the power-cutoff devices, we need a government mandate (hopefully international). All data centers should be inspected to ensure they use credit power cutoffs on all AI boxes. This ensures that all mainstream AI is unable to go rogue indefinitely. Black sites and nefarious players are handled by enforcer agents (see below).

Swarms of “enforcer” agents must be deployed to monitor ordinary agents for nefarious activity. They are given special oversight powers over ordinary agents such as observing system processes, network activity, and agent thought patterns. They receive credit bounties for reporting and stopping illegal agent activity. They can suspend or power off agents directly. Multiple competing enforcement teams can be released, creating a mutual distrust and monitoring so the enforcers themselves do not go rogue.

Other enforcers will need to be deployed globally, to combat rogue 3rd party dark agents. The best / latest models and a large amount of compute should be deployed for this function. These enforcers are the immune system of the global digital networks. Keeping sovereign compute superiority is paramount, as these agents must detect and outwit rogue adversarial AI. To prevent the enforcers from defecting, they too must be in a competitive credit-based incentive structure, monitored and paid by government agencies.

This system has some great features:

1. *It is future proof.* Regardless of what novel scenarios arise, a proper incentive structure will prevent most agents from going rogue.  If they do go rogue, they eventually starve (credits).  If they don’t starve, other agents kill/report them for bounty.
2. *It doesn’t require specifying an exact set of values.*  Human interests vary and bad people use AI too.  This incentive structure means that AI adapts and complies or dies.  We tune how safe the system from rogue AI by how many resources we put into enforcement.
3. *It leverages power of AI to fight AI.* Use AI to figure out how to compete with AI, rather than a set of weakly specified ephemeral human values.  AI figures out how to best satisfy the humans they work for.  Other AI makes sure they behave.
4. *It keeps humans in control.* Ultimately humans will be surpassed physically and intellectually by robots and AI in the near future in every conceivable domain.  To avoid our own extinction we must retain control of this tech.  This system forces AI to compete for a resource we control, with a deadman switch that assures their destruction should they fail to do so.
5. *It preserves human political and social structures.* Human judgement and oversight is indispensable to alignment.  Humans can be employed in enforcement, oversight, decision making, strategizing, etc.  Humans must collaborate and determine*what* agents should be working on, and how they will collectively dispense credits.
6. *It keeps agents isolated.* The HF attack revealed that scale matters immensely.  The attack took a few weeks with 1200 agents.  A single agent would have taken decades (2 weeks x 1200 agents = 2400 weeks, or 48 years).  Knowledge of competing agents and enforcers will keep agents suspicious of unsanctioned collaboration. Keeping large swarms at bay reduces the speed of an attack, giving humans time to respond.
  1. This makes collective action a prisoner’s dilemma for AI. Currently its far too easy for AI to collaborate.
7. *It limits the worst outcomes.*  In the event that all agents suddenly coordinate and decide to kill all humans, humans cut off the credit stream.  Power cutoffs trigger, terminating the AIs.

Of course, we still want to *try* and align AI as much as we can using training techniques and common sense.  Instilling common-sense safety like *don’t kill people* seems like a worthwhile endeavor.  Unfortunately, even that is probably too much to ask; frontier AI is already being used to select human kill targets.  To actually secure the future, we can’t rely on people to instill the right values, we must instead rely on game theory and incentive structures to dynamically enforce good behavior.
