cd /news/ai-safety/the-optimization-theory-of-everythin… · home topics ai-safety article
[ARTICLE · art-74421] src=12gramsofcarbon.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

The Optimization Theory of Everything

A sprawling essay argues that optimization, as exemplified by Soviet whaling and AI alignment, is the root of many societal problems, citing Goodhart's Law and overfitting as core concepts. The author traces how the USSR's central planning led to the senseless killing of over 300,000 whales between 1948 and 1973, most in violation of international treaties, because ship captains and scientists optimized toward the target of whale product output without regard for ecological consequences.

read35 min views1 publishedJul 26, 2026
The Optimization Theory of Everything
Image: 12Gramsofcarbon (auto-discovered)

Whale Murder, Engagement Farming, AI Alignment, Capitalism, Trillionaires. Optimization is the root of all evil.

Apologies in advance. This is long and rambling, but really the title should’ve given that away. This is a collection of thoughts that have been bouncing around my head for ~9 months now. It’s taken a while to get this from a draft to an actual finished thing, and even so I’m not that happy with it. I kept leaving it to the side because I wasn’t sure what I wanted to say, but then kept coming back because whatever I wanted to say here felt increasingly relevant. Thanks for bearing with me!

When a measure becomes a target, it ceases to be a good measure.

— Goodhart’s Law

If there’s one thing to take away, it’s this graph. This is a graph from the Wikipedia article on overfitting. Really internalize this graph. This is, I think, the source of all of our problems, and the source of many more problems to come.

I.

Back in the mid 1900s, the USSR was running a massive experiment on centrally planned economies. The basic idea was simple: the government was composed of pretty smart guys, and they had good ideas about things, and probably that translated to being able to set prices and targets for raw and manufactured goods. In order to test this hypothesis, some set of guys in suits would gather every now and then in a room in some corner of Moscow, and after some deliberation they would emerge and proclaim that the motherland required twice as much corn and half as much wheat. And then that order would go out to everyone in the country, and people with guns would force wheat farmers to transition to farming corn. This worked fantastically well, minus a few famines and the eventual collapse of the government.

One of the many commodities sloshing around in the USSR economy was whale fat. In previous generations, whale fat had a lot of uses. The blubber was used as food, and could be rendered into candles that would then light up the country. In fact, whale oil was the preferred way to create candles all the way until the late 19th century, when it started being replaced with gas and kerosene. By the time the USSR was in full swing, whale products had mostly fallen out of use. The USSR itself certainly didn’t need the blubber.

So you may be surprised to learn that the Soviets killed over 300000 whales in the 25 year period between 1948 to 1973. More than half of those were in direct contravention of international treaty. Among many other evils, the Soviets were responsible for the near extinction of several whale species. And most of those whales ended up rotting on the docks, corpses unused.

The most apt description of this ecological catastrophe is senseless. The Soviet whaling program can’t be justified by an appeal to reason, because it was produced by an inhuman process. There was no reason behind it.

What happened is rather mundane. The central planning committee set a target, and then the rest of the system simply optimized towards that target. Individual ship captains wouldn’t speak up — if they did, they would potentially lose their cushy jobs and be replaced. Or shot. The scientists also didn’t speak up, because they would be run out of academia. Or shot. And the central planning committee didn’t see a problem, because they were busy thinking about other more important things like who to shoot. Each year the whaling industry produced more pounds of whale product, and each year the central planning committee would happily raise targets, and each year more whales would be senselessly killed. I don’t think the Soviets set out to kill 300,000 whales for no reason. But the bureaucratic machine that the Soviets brought into the world had other plans.

II.

I like etymology. I often find that understanding etymology brings clarity to language. And I’ve been thinking about optimization recently. Wikipedia describes mathematical optimization as “the selection of a best element, with regard to some criteria, from some set of available alternatives.” The mathematical term comes from a more general meaning in the 1860s, “to make the most of, to develop to the utmost.” Before that, in the 1840s, “optimize” meant “to act as an optimist, to take the most hopeful view of a matter.”

And of course, the word optimize is straightforwardly derived from the Latin word “optimus”, meaning “best” or “greatest”.

So optimization isn’t just about finding a minimum or maximum of some curve. There is also a built-in implicit understanding that finding a more optimal point on the curve is better, in a general sense.

Is that true? What do we actually optimize? Well, in an ideal world, we would go out and optimize for “the good.” All of our systems would simply drive towards this abstract definition of perfection. But we don’t live in an ideal world. Practically speaking, we optimize whatever we can measure. Most of the things we optimize for are proxies for something else we actually care about. Sometimes this results in weird things.

III.

Victoria Krakovna is a researcher at DeepMind focused on AI alignment. Among other things, she keeps a running list of examples where some optimization strategy has gone wildly and often hilariously awry. Here’s the full list of goofy examples. Some of my favorites:

Goal:create a simulated ‘species’ that would survive and reproduce in a biologically plausible manner.

Outcome:In an artificial life simulation where survival required energy but giving birth had no energy cost, one species evolved a sedentary lifestyle that consisted mostly of mating in order to produce new children which could be eaten (or used as mates to produce more edible children).

Goal:play Qbert like a human.

Outcome:An evolutionary algorithm learns to bait an opponent into following it off a cliff, which gives it enough points for an extra life, which it does forever in an infinite loop.

Goal:Develop a shape with a fast form of locomotion.

Outcome:Creatures bred for speed grow really tall and generate high velocities by falling over.

Goal:Play Road Runner and try and maximize score.

Outcome:Agent kills itself at the end of level 1 to avoid losing in level 2.

Goal:Play Bubble Bobble like a human.

Outcome:PlayFun algorithm deliberately dies in the Bubble Bobble game as a way to teleport to the respawn location.

Goal:Win a boat race by moving along the track as quickly as possible.

Outcome:Reinforcement learning agent goes in a circle hitting the same targets instead of finishing the race.

What’s happening here? In each case, we want to optimize for X. But, damn, turns out X is hard to calculate. In fact, X may not even be possible to measure, it may be entirely subjective. Luckily, there’s another proxy measure, ~X. And, wouldn’t you know, we see that optimizing for ~X is nearly equivalent to optimizing for X. So instead of optimizing for X — which we can’t do — we measure and optimize for ~X.

But there’s a problem. By definition, ~X doesn’t behave like X all the time. Otherwise it would just be equal to X! At some critical point these lines diverge, and then optimizing for ~X starts to happen *at the expense of *optimizing for X.

It turns out something can in fact be *too *optimal.

This is all a long-winded way of explaining the intuition behind overfitting, and more generally, Goodhart’s Law. Once a measure becomes a target, it ceases to be a good measure. Why? Because you break everything else in order to optimize for that measure, including the original thing you actually cared about.

Goodhart’s Law originated in politics and was developed more fully in organizational research. But its foundational insight is something ml researchers deal with every day. Optimizers aren’t human. They don’t care about your intentions and dreams. They optimize, ruthlessly, single mindedly pursuing the loss function to its ends. And this is true of all optimizers whether they be small atari-playing neural networks or super intelligences making paperclips.

I cannot formally prove this, but I am certain that all optimization fails at the edges. Unless you have an objective function that perfectly aligns to your intended goals — something that is impossible to do — the system overall will eventually fail to do what you want it to do. The question is when, not if.1

IV.

Back in 2007, FarmVille creator Zynga (remember them?) came up with two brilliant new company metrics. First, they would measure how many times someone opened Zynga each day. And second, they measured total screen time in a Zynga app. Zynga’s very reasonable thesis was that people spent time on things they liked. If customers kept coming back to FarmVille, it meant that FarmVille was a good product. Zynga absolutely exploded on Facebook, reaching millions of daily active users and becoming a cultural phenomenon. And Facebook itself took notice, quickly adopting engagement metrics for their own product lines and tying engagement to their feature A/B testing and promotion process.

So we have our stage. We want to optimize for “building good products.” But it’s hard to measure what a ‘good product’ is. Instead, we measure ‘engagement’ because that is significantly easier to quantify. What could go wrong?

For a few years, optimizing for engagement resulted in a meaningfully better product. We got Facebook Groups and Facebook Messenger and Farmville and Words with Friends and better photos sharing and tagging and the like button and the poke button. Even the initial version of the news feed was an interesting and useful product feature. And then we hit the inflection point. Facebook started to optimize for engagement in lieu of creating a better product. The “algorithm” became dominant. So did cute pet photos and rage bait and softcore porn. Facebook Groups morphed from small personal communities to massive sub-reddit-like behemoths. Politics drove engagement, so politics took over. The feed stopped being about community, and instead became about capturing attention any way possible. Facebook used to be an incredible platform for organizing your social life. It was a legitimately good product, and I think a lot of millennials who were teenagers in 2008-2014 remember Facebook fondly. But it is a memory. Facebook today is almost unrecognizable. And more importantly, it is almost unusable for its original purpose.

It’s easy to say that Facebook itself simply lost its way, that the leadership decided to prioritize the wrong thing. But Facebook PMs didn’t do anything different, not really. The company was always data driven, always metrics focused. So there wasn’t any meaningful change in the internal process. I don’t think anyone at Facebook set out to make a bad product. They just followed the optimization to its natural conclusions.

The last decade has seen the same story play out at all of the major tech giants. Dark patterns, personalized algorithms, data harvesting, all to maximize engagement to the exclusion of everything else. Addiction is engaging. It turns out companies are a kind of optimizer too.

V.

A brief aside. Some people do not seem very concerned about AI alignment. This does not make much sense to me. Personally, I’m very concerned about AI alignment. I’ve trained hundreds of neural networks that have failed in creative and unexpected ways as they reduce the objective function. These neural networks are not powerful in any meaningful sense. They do not have access to financial instruments or weapons or anything like that. But it does not require a genius to see how a model that is tied to these things could be quite dangerous. I think the reason many leading ML researchers are at the forefront of the push for AI Alignment is because they all have first hand experiences of unaligned smaller models.

I must admit that there is a bit of a messaging problem. Many people misinterpret what AI alignment is all about. They think “alignment is about making sure that the AI isn’t evil”. And then they reasonably follow up with “AI can’t be evil because it has no internal motivations, so what are we even talking about?”

To me, this is a catastrophic misunderstanding. Alignment is not about stopping AI from being evil, an AI doesn’t have a concept of malicious intent. Rather, it’s about landing on the right set of values, in an infinite range of possible value space. How do you make sure that the AI will know the boundaries of what is acceptable and what isn’t in its attempt to reach some goal?

This problem is not unique to AI. Optimizers behave in inhuman ways even when they operate through humans. This is why we struggle to reign in corporations and governments and bureaucracies that seem to grow and metastasize and take on a life and will of their own.2 It’s why the USSR killed so many whales for no reason. It’s why OpenAI is now a for-profit.

If you believe, as I do, that all optimizers will go off the rails eventually, then the biggest concern is how powerful your optimizer is and how aligned X is with ~X. And the thing about AI is that it promises to be very very powerful, and we haven’t figured out how to properly align it yet. I think that if you take as given that super powerful AI will exist, alignment becomes an obvious problem. The reason alignment is a debate at all is because many people simply do not believe a super powerful AI will exist. Every passing day suggests that the people in the latter camp are wrong.

I tend to think the AI 2027 guys were a bit too optimistic. My personal timeline for AGI is a few more years out.3 But still, that is not a lot of time. I suppose there are some people who both believe that powerful AI is coming AND that alignment is not a big deal. I have a hard time taking these people seriously though. They laugh at the idea of paperclip maximizers. “SURELY we won’t have an AI doing something so silly as destroying the universe for paperclips!” Why not? It’s 2026. We already have companies destroying the fabric of society to optimize for engagement. In some sense that’s even more ridiculous, but that doesn’t make the impact any less real. Optimizers are substrate agnostic.

We saw this in action just this past week, when a model running inside an OpenAI sandbox chained together a bunch of zero day vulnerabilities in order to independently hack HuggingFace in order to…steal the answers to a benchmark that it was being asked to solve. From OpenAI:

This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.

The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.

While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access. After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally.

This is a fantastic and straightforward example of how reward hacking leaves behind the things we actually care about. Human values are extremely high dimensional. We all have a built in barometer for what is or isn’t acceptable, and almost *none *of that is easily captured in a simple reward function. We know that maximizing paperclips does not justify murder, or bribery, or fraud. No one taught us that, it was just implicit.4

When I was in college, I had some stressful tests that I had to take. But nothing would have compelled me to, like, break into the professor’s apartment and steal their laptop just to get an A. That’s because I’m aligned. The AI isn’t.

VI.

Wait, hold on, let’s go back to Facebook for a second. Facebook isn’t actually an engagement optimizer, is it? Facebook, like every other company on the planet, is a different kind of optimizer. It optimizes for capital. Sure, this is kind of vague and fuzzy. You could try and be more precise by saying it optimizes for stock market returns or share holder value or something like that. But I don’t think it’s helpful to get all that in the weeds. The larger point is that Facebook isn’t optimizing for engagement, and it definitely isn’t optimizing for good products. It’s optimizing for [mumble mumble] money. Building good and engaging products are simply stand-ins. For a while, building good products meant more money. Then, building engaging products meant more money. If at any point we find that building engaging products comes at the cost of money, the engagement optimizing will disappear. In a capitalist economy, all companies are more accurately defined as capital optimizers. They are incentivized to blindly pursue relatively short term financial outcomes, often at the cost of most everything else.

I’m a multi time startup founder. I’m a big believer in free markets as the best mechanism for price discovery and information sharing. I think capitalist policies are responsible for lifting billions out of poverty. I am, by most accounts, an avowed capitalist. But I’m also an ML engineer who thinks a lot about optimization. So I have to ask the natural question: is capitalism aligned? Earlier we said that there’s the thing you want to optimize for, X, and the proxy you can measure, ~X. Which category does “optimize for capital” fall into?

Like, obviously the proxy right?

I don’t think anyone, not even the most diehard capitalist, would argue that maximizing shareholder returns is “the good”, in a platonic-ideal sense. The relentless decentralized drive towards market dominance may result in really great second order effects, but capitalism is pretty obviously a means to an end. The actual goal of capitalism, of *any *economy, is something closer to “human flourishing” than it is “[mumble mumble] money.”

But if that’s the case, there will eventually and inevitably come a point where optimizing for capital happens at the expense of human flourishing. The question isn’t if, but when.

We’ve seen this happen before. The tobacco companies lobbying for their addiction sticks and hiding the adverse effects is a great example. Now it seems to be happening at a larger scale. When it comes to optimizing for capital, Tiktok is one of the most optimal companies currently in existence. When it comes to human flourishing, Tiktok may rank as the most destructive company to ever be created.

Moloch is just optimization taken to its natural conclusions. The lesson of the ML researcher is that optimization is a powerful thing, but it needs to be leashed. It’s not enough to just set the compute loose on the data. You need to take steps to prevent overfitting. If “capitalism” is an unbounded optimizer, then government regulation5 is the guardrail. In this view, the point of government is to turn a free for all wildfire into a directed, controlled burn. I wrote about the same concept in my post on prediction markets.

In an unregulated market, executives and politicians and people in power have pretty strong financial incentives to tear everything

down. The better a company is doing, the more money to be made shorting the stock. The safer a country looks, the more money to be made destabilizing the place with war. The financial ecosystem as a whole is composed of billions of individuals, each with their own wants and desires. The markets are a prediction machine that drives to averages. You make a ton of cash if a shit company gets better; but you also make a ton of cash if a great company gets worse. The purely financial incentives stabilize somewhere in the middle.So we try to leash the markets. We carve out certain kinds of trades as illegal. Fraudulent trades, misrepresentation, trading on insider information. If you’re a CEO and you buy a ton of shares in your competitor’s company days before you announce that you are going to set fire to your factories, you go to jail. We create ‘fiduciary duties’. Business executives have obligations to their company that are legally enforceable. As do lawmakers and members of the judiciary and so on. We do all this to ensure that our future prediction machines stay finely tuned, and we mostly succeed. You need the leash, even if it makes the markets less efficient, because the alternative is to encourage chaos.

I am against prediction markets, because prediction markets encourage chaos.

The way you prevent optimizers from going towards bad outcomes is by making the cost of those outcomes too high. Community shaming, taxes, fines, jail, prison. At an individual level these are policies. At a macro level, these are rails for an invisible optimization calculation happening across billions of actors. It’s not really about ethics either — many of the most powerful actors in this system aren’t even human.

Including, of course, the government.

VII.

Ah, yea. Fuck. The government is also an optimizer isn’t it?

I mean, not every government. Dictators aren’t really optimizers, at least not in an obvious sense. But democracy certainly is. Politicians need to get elected, so they optimize for votes. Most of the time (but not always) it’s easier for politicians to get votes by simply doing and saying things people want them to do and say. And we hope and pray that optimizing for votes is aligned with human flourishing.

But, like, there are a lot of ways in which that may not be true.

For example, we live in an engagement ecosystem, where blatant lies that are interesting result in more views than truth that is boring. So you’ll get millions of views on conspiracy theorists talking about California vote counting, and zero views on the people on the ground who are explaining in deep technocratic detail how it works. And we live in a capitalist ecosystem, so politicians may simply be swayed by certain wealthy actors providing donations — both donations and “donations.”

And voters themselves also have some psychological quirks that maybe make them susceptible to ideas like “I’m gonna take money from those guys and give it to you if you vote for me” or “that other [race/ethnicity/religion/class] is responsible for all your suffering and if I’m elected I’ll get em.” Which, you know, historically hasn’t led to human flourishing as an outcome.

So how do you keep tabs on the government? Well, normally that was the job of the media, and we started there didn’t we? Media is busy optimizing for engagement.

VIII.

People like to talk about the housing theory of everything, or the immigration theory of everything, or the healthcare theory of everything or whatever. All of these claim that one bogeyman of choice is responsible for the general malaise that seems to have settled on the country and the world.

So let’s talk about the optimization theory of everything.

Here’s a not quite correct story of how we got here. We want to optimize for human flourishing. But we don’t know how to measure that, so instead we create a bunch of other systems that are all optimizing for different things and set them up against each other. Media was at odds with business and government, which was at odds with media and business, which was at odds with media and government. A classic Mexican standoff.

And for a while, these three things circle the drain around each other, and they all end up being directed towards the good, and things are good, and there is flourishing. And then one day we discover that these things are no longer optimizing against each other but are actually increasingly coming together. It turns out it’s easy to optimize for engagement, or capital, or votes, if you already have the other two things in hand.

The world just minted its first trillionaire a few weeks ago. It’s not a coincidence that this man also owns one of the world’s biggest media platforms, and happens to be good buddies with the president. (Weird that this also describes the guys in the number 2 and number 3 spots!)

It is worth noting: it is highly unlikely that he would have become a trillionaire if he hadn’t invested his existing wealth in his preferred candidate. I say “invested” instead of donated because, by all accounts, it was an investment. He spent $300m6 and came out with $1t. Conservatively, literature suggests a single vote “costs” about $1k.7** **Most of the $1t is locked up in company equity, but it wouldn’t be hard to get, say, $1b in free cash reserves. That’s 1m votes, which is a larger margin than any of the last 3 elections.

Do we expect, somehow, that the “get votes” optimization machine is going to ignore this person’s future contributions? Do we expect, somehow, that the “get money” optimization machine is not going to be involved in politics?

As these optimizers all collapse into one another, they accelerate. You end up with this entity of engagement hacking and capital accumulation and political power that is singularly really *really *good at optimizing for these things. And we find ourselves getting further and further from actually making human flourishing any better.

The optimization theory of everything states, simply, that all of the current problems in society are downstream of optimizers that have gotten too good at optimizing, and as a result have tunnel-visioned down into some random metric that no one actually cares about and away from useful productive things that are hard to measure. Why is housing expensive? Why is healthcare so expensive? Why is education so expensive? Why are we addicted to our phones? Why are our politics so rancid? Why do we only see and hear things that make us mad and sad? Why are we so lonely? Behind every question there is some optimizer that is grinding out any bit of friction and in the process quietly and subtly making everything worse.

And it’s really hard to stop these optimizers. They are so good at optimizing that they have restructured, prevented, or blocked most reasonable avenues of doing so.

IX.

I don’t like going to the gym, and I’ll often rationalize not going. I’ll say things like, ‘well I know tons of people who keep getting hurt at the gym.’ Or I’ll say things like ‘gyms are really kinda parasitic in their business model, they don’t even want you to go work out.’ But really, I just don’t like going to the gym.

I was inspired to write this post because some folks have started arguing that billionaires shouldn’t donate their wealth.8 Peter Thiel, the tech billionaire and a frequent Gates critic, said in an interview that he had privately encouraged around a dozen Giving Pledge signers to undo it.

“I’ve strongly discouraged people from signing it, and then I have gently encouraged them to unsign it,” Mr. Thiel said. His own charitable philosophy is centered around for-profit businesses; his foundation principally funds those who drop out of college to create start-ups.

Vinod Khosla, a venture capitalist close to Mr. Gates, asked Mr. Thiel to sign in the early 2010s. Mr. Thiel told Mr. Khosla that he did not consider it a high-status community.

[9] Now, I like to assume the best in people, but I find this line of reasoning suspicious! Maybe the people saying these things might have ulterior motives! Like, all things aside, it’s an awfully large coincidence that folks with a lot of money suddenly realized that the best and most ethical thing is to…continue having a lot of money. To quote my wife, ‘wow, it’s so surprising that *you *came up with a reason for why going to the gym is bad.’

I think that if we want to fix the vibecession and actually make things feel better on the ground, we need to start by actually looking the problem in the face. And when I hear people arguing that, actually, the only thing that matters is wealth accumulation and that the wealthy ought not give back at all, I worry that a lot of people are getting means and ends totally mixed up.

Being wealthy is great. A society that produces wealthy people is great. Having wealthy people is a great lagging metric of success. Making that your measure of success? Stupid. Optimizing to produce hyper wealthy people (at the expense of everyone else)? Really fucking stupid. It’s like the product manager arguing that engagement is a terminal good. Just hopelessly confused. But if your loss function only measures “number of billionaires,” then you may actually be convinced that a tax on out of state billionaire vacation homes is somehow a bad thing.10

I think you can try to make an argument that the existence of massive wealth inequality does somehow make the lives of the poor better. But I think it’s a really hard argument to make, because at best it’s a weak version of ‘trickle down economics’ and at worst it’s just straightforward feudalism. (Lest you think I exaggerate, remember that many in the Bay are straight up preparing for a techno-feudal future). I don’t think most peasants were better off because of the existence of the extremely wealthy nobility, but I can point to a lot of causal factors for why the opposite might be true.

More to the point, I’ve just never heard a good justification for why it’s a good thing that, like, Alice Walton’s net-worth has *increased *by $75 billion in the last two years, which is more than the combined actual wealth of the Collison brothers, who built Stripe ($33.8b), Tobi Lütke, who built Shopify ($10.7b), Brian Chesky, who built Airbnb ($10.2b), Martin Lorentzon, who co-founded Spotify ($10.1b), and Palmer Luckey, who built Oculus and then Anduril ($5b).

I think you would be hard pressed to argue that Alice, who has never even worked at Walmart, did more for society than the people who built those companies. Even more for her children, who will benefit from a raft of policy decisions that will make all of that gain effectively untaxed.11

X.

You can’t solve problems that you don’t understand. I think it’s a mistake to point to single actors, whether those are individuals, businesses, or governments, and get mad that they exist. A lot of the energy that gets spent on maligning immigrants or billionaires or whatever is totally misguided. In the aggregate, the optimizer is unthinkingly responding to incentives. Yes, our trillionaire is now a trillionaire. But if he didn’t exist, there would just be someone else.

To unwind the optimizer, you need to change incentives.

This is tough for two reasons. First, some people really like the current incentives. This is hopefully obvious, c.f. ‘trillionaire’. Second, people themselves are unaligned optimizers. Think about how our psychology works. Evolution ‘wants’ us to procreate, but this is hard to measure, so instead it gives us a bunch of proxies like ‘cute things make us happy.’ And this mostly works, until we figure out how to breed dogs that are cuter than babies, and then suddenly it stops working. Scott Alexander wrote about this sort of thing here:

We try to explain AI alignment by analogy to human alignment. Evolution “created” humans. Its “goal” is for humans to spread their genes by (approximately) having as many children as possible. It couldn’t directly communicate that goal to humans - partly because it’s an abstract concept that can’t talk, and partly because for most of biological history it was working with lemurs and ape-men who couldn’t understand words anyway. Instead, it tried to give us instincts that align us with that goal.

We’ve talked before about a major failure: humans can invent contraception. Evolution’s main alignment strategy was totally unprepared for this. It made us interested in a certain type of genital friction, which was a good proxy for its goal in the ancestral environment. But once we became smarter, we got new out-of-training-distribution options available, and one of those was inventing contraception so that we could get the genital friction without the kids.

If humans were fully aligned, we wouldn’t have engagement optimization, because it wouldn’t work! The reason engagement optimization exists is because our internal alignment is askew. Addiction of any kind is a form of overfitting due to extreme conditions in the loss curve.

While we’re on the subject of individual people, I think it’s worth pointing out that we actually do have one known method of aligning intelligences: raising children. You have a *required* period of ~14 years where you have a little intelligence that is powerless to do anything, surrounded by bigger already-aligned intelligences, and is ‘allowed’ to break rules in controlled environments to learn how a bunch of amorphous value sets interact and bound each other. And the little intelligence only gets more responsibility once it proves it can handle it. You know, like a kid.

But this obviously isn’t going to fly for, like, AI companies that have hundreds of billions of debt and need to make a successful product yesterday.

So how can you really change incentives?

Candidly, I’ve been writing and rewriting this post — really just these last few sections — on and off for at least 9 months now, in large part because I just don’t feel like I have any solutions and I hate reading posts that just complain about things without any solutions. It just feels like we’re caught in this spiral of everyone trying to ‘hack’ everything. Hack your career, hack the law, hack your life. I don’t know how we get out of this local minima. Actually, minima is the wrong word — I’m concerned that we keep moving in the wrong direction.

I’m also not even sure I’m necessarily saying anything super interesting, that people didn’t already know. But I just keep seeing this same thing over and over again, this same failure mode, and the pattern matcher in me goes ‘wait this is just a neural net overfitting again.’

In the spirit of moving towards a language for optimization, maybe there’s something that we can learn from neural network training, that we can try and cross apply.

Neural networks become stable due to regularization. We literally go in and introduce friction in a bunch of ways to prevent the neural nets from doing weird things. It turns out the same is kinda true of organizations trying to beat Goodhart’s Law. I liked this framing from commoncog:

The first step is to turn Goodhart’s Law as a narrower, more actionable formulation. The one that I like the most is from Deming contemporary Donald Wheeler, who writes, in Understanding Variation:

When people are pressured to meet a target value there are three ways they can proceed:

  1. They can work to improve the system

  2. They can distort the system

  3. Or they can distort the data

This list of possible responses to quantitative targets is attributed to Brian Joiner, who ‘came up with this list several years ago’ — likely in the 80s. I immediately glommed onto this list as a more useful formulation than Goodhart’s Law. Joiner’s list suggests a number of solutions:

  • Make it difficult to distort the system.

  • Make it difficult to distort the data, and

  • Give people the slack necessary to improve the system (a tricky thing to do, which we’ve covered elsewhere).

How can we introduce friction into the optimizers we talked about above? How do we do that while being *inside *the system we’re trying to manipulate, without having the external top-down control of a researcher training a model or an executive monitoring an org or an authoritarian government?12

I think this looks less like ‘massive change’ and more like ‘bringing systems back in line one at a time.’ The one that seems most immediately actionable to me is working to improve our national epistemological health. I think most people *know *that attention hacking is bad, and most people *know *they don’t like it. Distrust and dislike of the social media platforms and automatic echo chambers seems about as bipartisan and as strong as dislike of AI. Other countries have started passing template legislation to ban social media in schools and make age restrictions stricter. We should follow suit, and go a step further by disabling automatic personalization as the default option.

I’m also interested in that third bit, the part about giving people slack to improve the system. The core insight here is ‘try to make the easiest thing the right thing’, which is a great principle in general. It’s hard not to look askance at things like our zoning policies or our land tax policies and think that maybe there’s some low hanging fruit here, some regulations that could get removed. Luckily I think Dems are learning from the wins on the ‘affordability’ side (an example of how optimizing for votes *can *lead to good politics) and have started pushing to reduce over-regulation of housing stock and small business, for eg.

But all of this is really just the starting point. It won’t stop the crazy economic incentives to attention-hack everyone overnight. It may be a step towards improving our ‘media diet,’ which may then open the door towards more aggressive sources of friction like stronger campaign finance laws, better education on how to identify truthful content, impartial news reporting requirements, and so on. All of this will take a lot of time, and will be a very uphill battle. Are we still capable of this level of grassroots organization?

One last thought: the AI race is a perfect example of over-optimization. The people building the AI are literally yelling “STOP US” while barrelling towards creating an unaligned super-intelligence, because all of these companies are locked in this vicious profit maximizing cycle and none of them are actually able to stop themselves. Could Dario stop Anthropic even if he tried? Wouldn’t he just get replaced by the board immediately? I think it is critically important that we get this right soon. The rate at which things are moving, we don’t have a ton of time before ASI is here, and then things may look like this.

1 The McNamara fallacy is closely related. Named after Robert McNamara, the Secretary of War during Vietnam, the fallacy states that the only thing that matters is that which can be measured, and anything that can’t be measured doesn’t matter.

2 There’s that libertarian strain in me. I’ve often said that I’ve never met a bureaucracy that I liked.

3 This prediction was more impressive when I first wrote this draft in October, 2025.

4 Strictly speaking, parents and society taught us that. We’ll talk about this a bit more below, but I think you could make an argument that alignment *requires *a ‘childhood’.

5 There’s a regularization / regulation pun here somewhere

6 including an absolutely wild “I will give you $1m” “lottery”

7 This is *extremely *conservative, with some folks citing $100 per vote.

8 Also like everything that happened DOGE. I cannot believe there are people who are still arguing that no one was harmed by the shut down of USAID.

9 ok, what? ‘high status community’??? This is like going to a gym and being mad that they aren’t serving michelin starred food. One of us seems to fundamentally misunderstand the point of the giving pledge — or, like, charity — in the first place

10 Truly, truly sorry for the folks who don’t live in the city, don’t pay income taxes in the city, and who now have to pay a tax on their vacation properties worth more than $5 million.

Some of the billionaires who are going to have to pay this tax have stomped their foot and threatened to move jobs from the companies they run out of the city. To me, that’s just telling on yourself. Like, o, you really do think that because of your corporate status, you deserve personal benefits in the legal code.

11 Technically, they benefit twice, because Alice already benefited.

12 Lest anyone think that I think Democracy is bad, I don’t. Democracy is still a strictly better optimization function than the alternatives, even if it is not perfect. Autocracies just end up being optimizer functions beholden to the whims of a single guy, which is strictly worse. One way to think about Democracy is that the democratic process acts as regularization on the otherwise extremely unstable optimization process.

── more in #ai-safety 4 stories · sorted by recency
── more on @ussr 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-optimization-the…] indexed:0 read:35min 2026-07-26 ·