# Enough Reason to Act

> Source: <https://substack.norabble.com/p/enough-reason-to-act>
> Published: 2026-09-22 11:40:56+00:00

The recent AI news is a renewed push for a slowdown. [The spark was a resignation (Jacob Coxon)](https://x.com/hilbertspaess/status/2097476196791709843). Industry support is broader and deeper. The penetration into everyday news is deeper. The reaction from politicians is louder.

My interest is the reason for acting. I make the case for that in [The Case for Action](https://docs.google.com/document/d/1BS9ryr4MJ9_PG4wwwP1oz3DqPWmvQS0gfCH4o78v1GQ/edit?tab=t.0#heading=h.c4rhnyu9bp6m). But first, I address eight mistakes that are used as arguments against acting. Four of these are, [Mistakes of Detail](https://docs.google.com/document/d/1BS9ryr4MJ9_PG4wwwP1oz3DqPWmvQS0gfCH4o78v1GQ/edit?tab=t.0#heading=h.nzsk2lgdo120), false assumptions or mistaken facts. Four are [Mistakes of Form](https://docs.google.com/document/d/1BS9ryr4MJ9_PG4wwwP1oz3DqPWmvQS0gfCH4o78v1GQ/edit?tab=t.0#heading=h.y79v6scgj2os), modes of reasoning that fail.

The core of the news cycle focuses on loss of control, aka rogue AI. One of the turning points was Dario Amodei, CEO of Anthropic, publishing, [We Must Pace the Frontier](https://darioamodei.com/post/we-must-pace-the-frontier). In that, Amodei mentions several risks:

*the risk of [losing control of AI systems](https://www.anthropic.com/news/improving-alignment-security-efforts), [misuse of AI for cyberattacks and bioterrorism](https://www.anthropic.com/threat-intelligence-report-september-2026), and [serious economic disruption](https://www.anthropic.com/institute/econ-scenarios).*

So far, I’ve covered cybersecurity ([Every Reward Bends](https://substack.norabble.com/p/every-reward-bends), [Nobody Was Watching](https://substack.norabble.com/p/nobody-was-watching), [Security Can’t Wait](https://substack.norabble.com/p/security-cant-wait)). Partly that’s because my expertise is stronger there. Partly because it requires broad action across information technology, and I can be more persuasive there. The many small incremental activities lend themselves more to the type of influence I have today.

This is my first serious effort at discussing loss of control. The most pressing aspect is that the important steps take time to arrange. They are difficult, and become more difficult if rushed. The direct actors are fewer and more centralized, but political support is necessary. That requires broad understanding that takes time to reach.

The timeline is less definitive. The next 12-18 months will be bumpy in cybersecurity. Loss of control, if it does emerge, might be farther away. But there is a chance it’s equally close. Warning signals can compel action, but they can also compel overreaction. To avoid that, we want to see broad awareness and creation of systems that facilitate coordination that mature. If we have time, they’ll mature smoothly, if we don’t we’ll be more prepared than otherwise.

## A Title More of Us Could Agree With

The focal point of loss of control as a narrative is the book *[If Anyone Builds It, Everyone Dies](https://ifanyonebuildsit.com/)*, a 2025 book from Eliezer Yudkowsky & Nate Soares. It’s an eye-catching title. If I were publishing the same book, I might accept the same title, because, you know, titles. But unlike Yudkowsky and Soares, I do not accept the literal meaning.

I can iterate this one clause at a time toward something with both my full agreement and broader acceptance. My first iteration adds uncertainty:

*If anyone builds it, then **(there is a chance)** everyone dies*

I consider five possibilities:

(a) we successfully control superintelligence by intentional alignment

(b) we accidentally create a benevolent AI overlord

(c) we accidentally create a non-benevolent AI overlord

(d) we accidentally create an AI that destroys human civilization

(e) we do (d), (c) or (b), but before losing all physical control manage to enact some hail-mary or emergency plan that actually works at destroying (b), (c) or (d).

It’s tempting to try and assign probabilities to each of these. They would have large error bars if I did. If it were the only iteration, I’d do that. But the following iterations decrease those probabilities’ relevance.

*If anyone builds it, then (there is a chance) everyone dies, **(but if we build it right, that chance is less)***

We have agency here. Those error bars will stay high, but we can push the probabilities down. You might still say, why take any risk, if we don’t build it, no risk.

*If anyone builds it **(and let’s be honest, someone is going to want to build it)**, then (there is a chance) everyone dies, (but if we build it right, that chance is less)*

The world is a big place, with a lot of people. Even if you and I believed the strongest version, (d), not everyone would, nor would we be able to convince them all.

Why would anyone still build it? Well, because both the accidental benevolent overlord (b) and especially deliberate alignment (a) appeal. Some reject (b) on a value basis. In a strict interpretation, I wouldn’t. You might even argue a large part of humanity has hoped this is true, but in supernatural form. I would have a problem with those that object to (b) being forced into it though. But it’s pretty easy to imagine a group unconcerned with either of our views, and how can you fully object to (a)?

So, if there’s going to be a lot of people who want to try at (a) or (b), you could try and stop them. But again, the world is big, so you need very wide agreement for that to be practical. And when they do try, you’ll be stuck with whatever that attempt produces.

*If anyone builds it (and let’s be honest, someone is going to want to build it), then (there is a chance) everyone dies, (but if we build it right, that chance is less) **(so we should be part of building it, so we can do it right)***

But if (a) is possible, or you can discover if it is possible, you can try and do that. And if you trust yourself more than some other segment of humanity, you can try and make (a) more likely, or prove it’s impossible, which could convince everyone to stop trying.

Is it hubristic to trust yourself? Yes, or at least it flirts with it. Is it avoidable? Only if you trust someone else to build it better.

If every modest person defers, the hubristic take the wheel. So, the capable but modest must embrace some self-trust and hope they haven’t misjudged themselves.

Seeking power for power’s sake is bad, and it’s easy to lose self-control in that process, but seeking power for an outcome is also a necessity. You can distrust this a little, but you can’t distrust it absolutely.

*If anyone builds it (and let’s be honest, someone is going to want to build it), then (there is a chance) everyone dies, (but if we build it right, that chance is less) (so we should be part of building it, so we can do it right) **(but we also need to not be under pressure to do it fast, or even our good intentions will fail)***

When you look at those from labs calling for pacing the frontier, their views are closer to the iterated title. Soares and Yudkowsky may believe in the simple version, and you may disagree with that. But that wouldn’t give you a reason to disagree with the more qualified version. And people in the labs who agree with the qualified version are suggesting a slowdown as critical to finding a solution.

The control necessary to permanently block everyone is higher than adding extra time by going slower. Going slower gives more time to do it right. It gives more time to discover what’s possible. You may even discover that superintelligence is impossible, or much farther away. I find the first very unlikely, and the second is becoming less probable, but who knows?

We should take that full concept seriously. We still have to decide how much any individual believes this. We still have to decide which ones handle pressure the best. We still have to decide which ones are actually most competent at building something safe. And we still have to get the “someones” who aren’t part of this group to agree, or take away their ability to succeed (quickly).

I think we’re capable of coming to a solution here. But if we rush, out of complacency, or because we’re fighting over something else, we’re far more likely to fail.

## Mistakes of Detail

Before I make my case for action, I want to address some common mistakes. These are details where it’s easy to misunderstand the general case for action. I’ve seen these often enough to think there’s a chance you might also make them. I do not know you will, but if by naming them I give you a better chance, it’s valuable.

### Presume we all have to die

A common mistake is to say, I don’t see how AIs could kill us all. By the time they are half done the computers will go off and they’ll have to stop. Or something like that. This objection is irrelevant. We don’t want any large group of deaths.

I’m not against asking the question, because you can learn from it, and what you learn can be used in places you didn’t anticipate. If you ask the question, I’d point out (a) once civilization starts collapsing our chance of killing ourselves, always a risk, goes up, (b) you have not, and will not, think through every possibility. It shouldn’t be necessary to explain how a superintelligent system is going to escape from its physical limits after it escapes from our virtual limits.

There will be possibilities none of us can anticipate. Using the worst case to dismiss all the lesser but still bad cases is not an effective form of reasoning.

### Presume false competition

Loss of control is one topic. But it’s not the only one. Dario mentioned four. One mistake I’ve seen is thinking there’s a competition between them. Some people try to spin that as a reason to ignore loss of control.

If you ignore probability and timelines, loss of control is the most significant. Cybersecurity is the most time sensitive and requires the broadest set of actors to respond. Bioterrorism requires DNA labs, a very special kind of actor, to respond, and good deployment and training practices by AI labs. Economic disruption hasn’t really emerged, and it’s the one part I disagree with Dario.

The actors who are most pressured to respond to all of these at once are the AI labs themselves. In many cases there are coordination challenges with incompressible durations, no matter how much effort is expended. Global coordination of pacing might accelerate with shared priority, but there’s no unilateral shortcut.

Cybersecurity does fall into a unique slot: unmitigated, its effects land in a year. Rogue AI has a broader timeline. We don’t know when or if capabilities increase to the level that makes it possible. It’s a good bet they can, but not in the next 12 months.

### Presume cybersecurity and loss-of-control are the same

A pattern I’ve seen goes like this: Cybersecurity risks could be managed by diligence. I can understand how those risks materialize. I think I can apply that to loss-of-control.

It’s great you’re thinking of cybersecurity. It’s important. Cybersecurity needs lots of small incremental changes, starting immediately. We need a unified effort and commitment to doing that. We need great tools to make it efficient. If we do nothing, we could end up with catastrophes.

But the [loss-of-control risk takes place in a different world where things have changed](https://x.com/yishan/status/2100112647811408291). We don’t understand that world yet. It might be that this different world is harder to enter, and even if we try it could be a long way away. But I see no reason to be confident of that. We’re close enough that we’re learning some things. Cybersecurity applies to this world, but it’s also going to be radically changed.

So no, cybersecurity isn’t enough as a solution. I think it’s incredibly important. I think we can and will respond. If we didn’t we’d have chaos that would make every other bad outcome a bit more likely. If we really focus over the next 12-18 months, I expect we’ll enter a more stable period that’s more secure and better managed.

But it’s not one challenge. Being able to handle the cybersecurity challenges of the next 12-18 months is not the same as being able to handle loss-of-control. Good cybersecurity could avoid a few mistakes here. But it’s not a full response and can’t substitute for one. Partly that’s because handling cybersecurity is going to depend on having access to more capable models than the attackers. If the model is the attacker, and it’s the most capable model, that advantage is gone.

### Those guys are weird

I have to thank Matthew Yglesias for highlighting this trend, in [A simple plan to save the world from rogue AI](https://www.slowboring.com/p/a-simple-plan-to-save-the-world-from). Once mentioned, it’s impossible not to see. There are eccentric people, and eccentric opinions. Being an eccentric person doesn’t make your opinion invalid. A lot of uninformed dismissals take the form of attacks on Effective Altruists, often bringing up Sam Bankman-Fried. Yudkowsky has strong opinions and does eccentric things. The concept of being analytical about altruism, vs. doing what looks good / feels good, doesn’t seem like it should be eccentric, but through history it turns out it is. And yes, if you drop into San Francisco, you’ll find some eccentric lifestyle choices. But for San Francisco, if I know the history, that’s actually more typical than atypical. Either way, none of it is a successful argument about risks.

More importantly, the existence of an eccentric opinion doesn’t make the milder, more defensible version of that wrong. The milder, more defensible version of loss-of-control is still worth paying attention to. [10% chance of end-game](https://x.com/AIImpacts/status/2099581080223576403/photo/1) is pretty important. Strange as it is, it can require balancing, but dismissing it entirely would be a mistake.

## Mistakes of Form

The next set of mistakes are about the form of reasoning. They are more complex as a result. What’s ultimately most concerning about these, is that they are used to cut short a search for understanding. The first four fit a learning experience, but the next four, cut that short, substituting a dismissal. If accepted they lead to a stark choice, of no action, or an overreaction as the only remaining choice.

Some are patterns that reappear enough to be well-known. This is not, however, an exhaustive list. I call out these mistakes specifically because I’ve seen them used to shut down best-effort conversations.

### Reasoning from cynicism

Reasoning from cynicism is not a sound mode of reasoning. As a tool to identify concerns, it’s perfectly functional. But it’s not a complete form of reasoning. For any given situation, if you use cynicism as a form of reasoning, you’ll arrive at different conclusions based on which actor you target your cynicism at. And many of those conclusions will be directly contradictory. The truth may not be represented by any of them.

I’ll hold out Matt Stoller as an example, telling us to [Stop Panicking About AI](https://www.thebignewsletter.com/p/is-artificial-intelligence-going). This isn’t a generalized attack on Matt. He serves as an example because he made mistakes, used cynicism as a component of his reasoning, **and** got a lot of attention.

He knows he’s not an expert on this topic, yet is willing to make fact-free proclamations like “Agentic AI is dangerous, but we can improve the products to limit damage.“ How does he know that? His fix has no clear line back to safety.

I have no problem with Matt adding his view about copyrights. Dismissing the views of the better informed is unnecessary for that. But the post has over 600 likes. Is it his pre-existing audience? Is it creators defending their turf?

Those are the hopeful answers, because the post is not a good discussion of AI. He competently recounts the events that are available elsewhere, but the analysis doesn’t engage with the actual topic. As a response it’s slop, latching onto a live story to re-promote an existing point of view.

One of the ironies is that his prescribed fix is for us to abandon a “nihilistic refusal to govern for the public interest”. I’d be happy to. The irony is that this nihilism is closely related to the cynicism we (the public) have applied to politics. I sit in an uncomfortable middle, thinking mass-cynicism of the concept of government is a mistake, but sharing the same feeling about mass-cynicism of tech companies. You have to be willing to acknowledge the possibility of a good politician or a good company for there to be any room for them to succeed.

### Reasoning from markets

Markets are useful tools, but at best they process information, not create it. So when predictions about impacts from AI are made, it’d be a mistake to use a lack of market signals to dismiss them.

I read Marginal Revolution regularly. I’ve had disagreements with Tyler Cowen or Alex Tabarrok, but generally found the starting argument informative. But, recently Tyler has been suggesting AI and technology experts are making predictions unscientifically.

In [“Act like it is science”](https://marginalrevolution.com/marginalrevolution/2026/08/act-like-it-is-science.html) he asks:

*1. Your outline of how, say two or three years from now, we might estimate the additional cybersecurity costs from the AI break-ins.  I am convinced that number is not zero, but give me your method please.  If it helps, here is [an estimate of past cybersecurity costs](https://users.ox.ac.uk/~econ0628/Cyber_Risk.pdf), done by top economists.*

*2. Your current numerical estimate of what those costs might end up being, of course to be tested against what actually happens over time.  Obviously, you can do this for a few different regulatory/safety scenarios.  This is one simple way to prove yourself largely correct, albeit with a lag.  (NB: you do not have to take this as a substitute for your preferred safety measures.  But please try to be specific in your predictions.)*

*3. A list of what stocks or other assets you have shorted, now.  Obviously if your answer to #2 is sufficiently low, you could answer here zero, as I would do.  I expect costs, but not so high that we cannot muddle through and have expected positive stock returns.*

He’s since been [doubling](https://marginalrevolution.com/marginalrevolution/2026/09/numbers-numbers-numbers.html), [trebling](https://marginalrevolution.com/marginalrevolution/2026/09/sentences-to-ponder-140.html) and [quadrupling](https://marginalrevolution.com/marginalrevolution/2026/09/callum-williams-on-cybersecurity-prices.html) down, posting cybersecurity stock price charts and saying that with them “we are getting somewhere concrete and scientific rather than just scare stories.”

(1) sounds reasonable. But Tyler is responding to events that stirred both loss of control and cybersecurity concerns, and conflating the two.

Applying that request to the Soares and Yudkowsky version, that everyone dies, is absurd. The cost of everyone dying, crudely, is 8 billion x $14M = $112 quadrillion. That number is useless for any argument about what to do next. Historical data does not provide a probability, and we would be surprised to ever accumulate enough data to retroactively apply.

This problem of having no historical precedent doesn’t go away. Instead of one, all encompassing number, you start having many scenarios, each with an outcome (e.g. 1% die, so only $1 quadrillion). But each is just as useless. The concept is more useful than the number.

The second problem emerges in (2) and (3). We’re only ever going to get one version of “what actually happens over time”. But anyone who takes (1) seriously isn’t going to wait around. They are trying to change what actually happens. And (3) is a trust test, indicating Tyler is not going to take answers to (1) seriously, unless you put your money where your mouth is.

The [preparedness paradox](https://en.wikipedia.org/wiki/Preparedness_paradox) disconnects (1) and (3). Tyler can’t be fully unaware of this as in (2) he allows “*do this for a few different regulatory/safety scenarios”*. But you can’t pull those multiple scenarios forward to (3) which demands “your answer to #2” as a single number, without locking in probabilities that people will/won’t act.

If you follow his posts after, it’s clear that he’s not avoiding that trap, and (1) is just an opening salvo to draw you using insurance rates and stock prices instead of stories about events without precedent. Probably this is unintentional, and just derives from a habit of reasoning from markets in unfamiliar fields.

I do see a trend in Tyler’s posts toward [searching for evidence](https://marginalrevolution.com/marginalrevolution/2026/09/the-economics-of-cyber-risk.html), rather than dismissing evidence. But if he wants to do so, he needs to disclaim the original rather than just quietly walk away from it. This would help avoid others being drawn into the trap. He must also be open to non-market based predictions.

#### Shorts don’t tell you someone’s commitments to action

Shorts are an unfair test of credibility. Think through your options after you spot a risk:

**A pure trade.** You short, wait for the disaster, and collect mega money. You don’t need to do anything to divert. At best, you send a cryptic signal. If someone decodes the signal, they can trade on it too without taking action.

**A warning without trades.** You explain the risk, persuade others, plans are made and action is taken. You send a clear signal, no decoding needed. Action is the intended outcome. The beneficiaries are those who would have been hurt.

These are the simple options, but we can combine them:

**A warning monetized.** You short, then explain the risk and persuade others. Those that believe the risk won’t be managed act by selling, causing the price to drop. You exit the short, taking profits. Plans are made, action is taken, the risk is fully mitigated, and prices return to normal. The people who would have been hurt still benefit from the warning, but now you do too. The sellers end up a bit worse off, though it’s worth considering who they are. If they held the stock until the disaster that’s realized in the **pure trade**, they probably would have been hurt in that way.

You’ll notice that in **a warning without trades**, I didn’t mention sellers. It’s not core to that story, but realistically, if they appear in **a warning monetized** they’d appear there too. But there’s a reason they might not appear in either:

**A warning where trades don’t pay.** You short, then explain the risk and persuade others. *Others expect plans to be made and action to be taken, and don’t sell, expecting those to result in opportunity loss.* *Prices don’t change.* Plans are made, actions are taken. Your short fails as there was no disaster, and no opportunity to exit the position.

This scenario is bad for you, and once you’ve disclosed, there’s little you can do about it, other than manipulating views on the probability of action. In reality, it’s never quite so dire for short sellers, as there’s always a mix of opinions. But timing the exit is important.

**A warning where opportunity is missed.** You short, then explain the risk and persuade others. Those that believe the risk won’t be managed act by selling, causing the price to drop. *You don’t exit the short.* Plans are made, action is taken, the risk is fully mitigated, and prices return to normal. *Your short fails as you missed the chance to exit.*

If the risk is unavoidable or only partially mitigatable, there is always some residual payoff. But if we’re talking about risks that disrupt society, the silver lining may be inadequate. You may experience just as bad effects from the spillover as any residual payoff. And, well, everyone else suffers, so if you care about that, you really want plans that fully mitigate.

If the risk has low spillover, your best bet in shorting is the **pure trade.** If it has modest or more spillover, your best bet is **a warning monetized,** exiting prices reflect less confidence in mitigation than you have. If however, you expect prices will never dip to that point, you should never short in the first place.

It’s a little more complicated, because as discussion about a risk occurs, you learn more about the probability of mitigation. But this in-between period isn’t important to the question of whether you really can reason from markets. The only thing you’re guaranteed to learn is how persuaded people are and how that affects their participation in mitigation. There’s no guarantee you’ll learn anything new about the risk itself. If you do, it will be in the form of words, not prices. It will come from a person who has reasoned through the scenario. That person’s views could end up in prices, but you’d have no way to separate it from the other views that go into prices, except by listening to their reasoning.

You should not, in any direct way, be updating your views about the merits of arguments for risk, based on stock prices. At best, changes in stock prices are a signal of a change in opinion. Whose opinion and why? You don’t know. Better to just go to the source.

#### Insurance rates are also indirect

In another attempt Tyler Cowen suggests historical market prices and cybersecurity insurance rates are sufficient evidence that there’s not much to worry about. He assumes the market has already priced in the effects, that it’s also simultaneously going to force action.

The first problem is if insurance rates only reflect the views of insurance actuaries. If they aren’t yet convinced of a risk, rates won’t yet have changed. If you want them convinced, you either have to persuade them, or hope they persuade themselves.

The second problem is insurance actuaries also have opinions about the probabilities of mitigation. If you persuade them, you’ll likely persuade others. Rates would come back to normal the more they are convinced that mitigation is likely. Rates only reflect unmitigated risks and only to the degree that it’s likely to remain unmitigated up until an actual event.

The first and second problems are both affected, but in opposite directions, by the capability level of the actuaries. If they are wise all-knowing oracles about risks, they’d be wise all-knowing oracles about mitigation. They’d also have a responsibility to explain their rate setting process, which would disclose the risks. That would, assuming they are mitigatable, lead to mitigating those risks. They’d only fail to lead to mitigation of a mitigable event if their disclosures weren’t trusted. That would be odd if you accepted them as all knowing oracles.

The third problem is assuming the structure of finance isn’t hiding the risk. Finance has lots of fine print, and the cul-de-sacs of limited liability can swallow a lot of risk if strategically placed. There is the counter-argument to this that if there was a weakness, someone would step in and offer a new type of insurance or finance instrument to span it. But that’s in theory. In practice we have evidence of how this can fail. And keep in mind, there is a reason this is my third, not first problem.

#### Markets accumulate insight, they don’t produce it or disseminate it

If you’re not asking the hard questions, place faith in markets over reason, you’re in a circular type of reasoning that is a lazy response. If no investor is willing to hear the stories, there’s no reason for prices to change. If you successfully persuade investors to ignore the story, because prices haven’t changed, they won’t change until it’s no longer a story.

A perfectly functional market might be described as a representation of collective reasoning. But not only are markets not perfectly functional, collective reasoning has to begin with individual reasoning. Relying on the market could be a lazy way to capture others’ reasoning. But any appropriate update has to start from an individual, and then spread by other independent reasoning, or by person to person making arguments that don’t reference a market. Lazy followers will eventually adopt the same view once non-followers update their views, but the followers’ information will be delayed.

### Demanding yes-no answers

I wouldn’t claim I’ve proven anything here. Proving we’d lose control, or proving we’ll fail to respond to cybersecurity isn’t my intent. I don’t know that it’s even possible. These types of topics don’t adapt well to yes/no answers. It’s true in many domains, and very true here. Life is largely in the grey zone, and pretending like it isn’t only gets you so far.

We have a problem, with a public, that at an aggregate level–aggregating both individuals and individual actions–is far too partial toward yes/no explanations. It’s a tendency that has rewired our media, our politicians, even our personal interactions. It’s so pervasive, we take it for granted in most cases. When someone does bring it up, it’s treated as quixotic. Maybe I am tilting at windmills to make that argument.

It’s possible that the approach, or views, of Soares and Yudkowsky are more effective than mine in this world where yes/no thinking is as common as it is. I had to provide a long explanation, and a set of qualifiers. My book title would never have made it on the bestsellers list.

Another problem with my approach is I’ve just insulted the reasoning systems of a lot of people. That’s going to seem elitist if anything does. But I do it because I think they can change. I believe in them, and any resemblance to insulting a person is incidental. I don’t want to critique you, I only want to critique a reasoning system that ends in bad results.

The thing about “inexorable forces” is that many of them aren’t. They are powerful, hard to resist forces. But you can resist. If you want to successfully resist, you better plant your feet more than 2 inches away from the cliff.

For things that are difficult and are going to take time, you have to start earlier, and often work concurrently, taking care of other in the moment issues while making progress on the harder ones. It’d be nice to be able to focus on one issue at a time, but that’s just not how life goes. You may ignore and delay some big difficult things, but if you do, it’s primarily because they just aren’t important enough.

Importance here has a bit less to do with probability than other places. Early effort costs less than late effort. You commit to the big effort at a time when probability is more certain. The cost of trying to make up for missing the early effort, goes up with delays. So if it might be important in the future, you want to invest in things that shorten that future timeline.

We already have missed some opportunities. The most significant one is to have government institutions that understand this stuff. They would be delegated to create a plan, in collaboration, and when confident in that plan, move forward. Right now we don’t have the careful drafting of plans, we have reactions.

### Letting forcing functions win

The non-benevolent, superintelligent, super powerful AI risk has two versions. One is the event itself, anchored to the moment control is lost. The other version is worldly pressures forcing us to handle that pivotal moment poorly.

There’s some version of this that will always be about doing things in the moment, while being in a fog where we aren’t entirely sure if a superintelligent AI exists, or how close it is. That version relies on momentary things like, monitoring deployments, securing sandboxes, running evaluations. These are the things that if there is a precipice and we approach its edge, we’d see the edge. Logic says we stop there and admire the view.

There is another version that is about doing things in the moment, knowing that logic could fail, and a crowd, wanting to look over the edge themselves, might push us off. That version is about creating verifiable compute, about discussing the edge with places like China, about having a lot more people who understand some of this. Those three, are actually in order of difficulty, from least to most.

There are two ways you can let forcing functions win. First, you can give up, and not fight. Second, you can ignore that they exist and end up dominated by them. You have to be aware they exist and willing to work against a deck that isn’t in your favor if you want to take the reins and choose a path of your own choosing.

## The Case for Action

So far, I’ve only made the case for action obliquely, referring to others’ arguments and explaining how there’s a consistency between working in AI, and wanting it to be paced. On the cybersecurity side, I’ve done that already, in [Why It Hasn’t Happened Yet](https://substack.norabble.com/p/why-it-hasnt-happened-yet).

On the loss of control side, I haven’t. One of the critiques you see applied to the Soares and Yudkowsky version is specifics about how AI would get out of control, start to harm us, stay out of control, and end up killing all of us. I could provide a story like that, but it would have a sci-fi tinge where you could poke holes at. It’s more useful to think about how lesser forms of those stories are very possible, and that the consequences before regaining control, and the costs to regain control, are quite problematic.

We could go back and forth debating these stories, and likely never have a full resolution. The defenders would have holes where they couldn’t promise defense, and the rogue AI story would have holes where some unexplained jump would be necessary, given current capabilities and what humans have currently wired AI into.

This sequence of stories has less to do with absolute answers than possibilities. Those possibilities then connect with trajectories. One trajectory is how future possibilities are emerging. A second trajectory is our preparations. And the third trajectory is how quickly we can change either of these trajectories.

The OpenAI HuggingFace event revealed important information about these trajectories. It showed that labs were being reckless. If that continues we’ll be less prepared for whatever comes. It showed we had no structure to reduce recklessness. If that continues, we’ll do less preparation.

We also have the issue that the general public remains mostly uninformed about how information technology, especially this technology, works. It’s been a blind spot in our culture that accurate depictions of software developers are missing from popular culture. With that understanding missing, the popular reaction to events will both over and under react, and fail to choose the best nuanced response.

The risks to cybersecurity were evident within the technology domain before the OpenAI HuggingFace event. But another change from that event was to provide a demonstration that did not require technology understanding. From this, the public became aware of capabilities they would otherwise have struggled to understand.

That’s a big hurdle to overcome, and it’s just the first one. Creating a US regulatory system would take time. Many people who don’t understand each other’s positions need to understand them. A lot of details that don’t have shared understanding need to be agreed to. And then that’s followed by an international system, which has an even wider gulf and has to somehow co-exist with a set of standing disagreements at that level.

All of this should advise starting now. That doesn’t demand an immediate overreaction. In some ways, an early start is a way to avoid overreactions later. More importantly, it makes underreaction less likely and less inevitable. With a late start, the choice of actions immediately available may be poor choices all around.

An event might happen where the only reasonable action is shutting things down. That might go as broad as huge numbers of servers, or shutting down parts of the internet. With such a stark choice, inaction might take hold and nothing gets done until things get worse.

### How bad would it be if we did nothing?

I think we will take action. I speak about it to make that more likely. I believe others do too. I believe all this will mitigate many risks. So the risks I mention next, those aren’t risks I think will happen. They are risks I think we’ll avoid. If we do the bare minimum we will end up with some ugly costs.

Existing cybersecurity depends on successful layering of imperfect layers. Attackers have to search through every layer and compose attacks to perform their malicious activity. AI enables the search and composition.

The capabilities there are such that if nothing had been done, we would be seeing widespread successful cyberattacks. Fortunately, while more is needed, much has been done. The infrastructure for implementing cyber defense has history and maturity. AI was used to find, fix and deploy many critical vulnerabilities. Additionally, all the leading AI labs took responsible actions to restrict attackers’ attempts to use AI for composition of attacks.

For cybersecurity, you need to worry about how failures compound. For financially motivated attackers, if the success rate of an attack doubled, not only would you see twice as many successful attacks from existing efforts, but since the efficiency increased, you’ll have more effort. Second to that, the tax on revenue and then profits would accelerate as more are sapped by attackers. If you start with 1% of revenue, then double to 2% when success increases, then double again to 4% when attackers are incentivized to invest more. An average profit margin of 8% would be cut in half by those effects.

That kind of effect would force a reorganization of many business operations. That forced reorganization would create costs in any situation, but occurring suddenly, it might be poorly managed and result in systemic failures.

The loss of control risks could be described two ways. We could talk about the cost of effects a rogue AI might take. And then we can talk about the costs of the reactions we’d need to regain control. This assumes we do regain control.

I think people are generally right when they suggest that a rogue AI, at the current capability level, would only do limited damage. The costs to shut it down might be trivial, though there are scenarios where something worm-like emerges that requires wide scale computing outages to respond to. Generally though, concerns about rogue AI don’t focus on current capabilities. They worry about our ability to slow down or stop capabilities progress once we commit to it.

Without anything like global coordination, there are significant loose-ends to controlling capabilities progress. Some additional capabilities might emerge from low-compute evolutions, like better harnesses.

The actual [list of bad things a rogue AI could do is quite long](https://thezvi.substack.com/i/216043508/a-specific-detailed-story-about-ai-killing-everyone-that-doesnt-sound-to-me-like-science-fiction), in the same way as the list of bad things a person could do. What’s the minimum cost you’d need to suggest to merit an AI safety program that redirects half of all training capacity toward safety? Two thirds?

[Plan A](https://ai-2040.com/) uses this mechanism, detailed in its appendix.

### What are we really giving up?

Absolute outcomes like “everyone dies” can justify absolutely any response. The only logical limit to a response that avoids that is a less costly response that also avoids the outcome. But I’ve conceded that the outcomes may be less than absolute, that those are outside chances.

We don’t need absolute outcomes to merit responses, but once they aren’t absolute, we start needing to consider, what are we giving up when we take actions to avoid those possibilities? Solutions don’t just compete with themselves, they compete with inaction.

We have many ways of evaluating what the costs are but we often summarize those through the actions of our economy. A well functioning economy is important. In many senses we treat it like a background we don’t need to understand, but really, we do.

I have cross-over experience here. In many ways, I started this Substack with more interest in economics than tech. My experience in tech is even deeper, so I’ve gotten pulled back into that. I think my cross-over has something semi-unique to offer.

What would happen economically if we “paced the frontier”? There is some group that reacts to this as if it would be a catastrophe. It’s actually a small group. But it’s a group worth responding to. There is no reason to expect that as a result. The AI capabilities delivered already are quite enormous. The economic opportunity available by simply adopting them will provide a lot of value for quite some time. Value from making people’s lives better, and value from markets continuing to function. I’m more concerned with the first, but also acknowledge how the first depends a fair bit on the second.

We might see a few less data centers built, but honestly, that might not change much anyhow. We can divide the usage of compute into three categories: inference, capabilities research and alignment research. Demand for capabilities research declines when pacing, but increases in alignment research would take its place.

Inference demand is driven by adoption and moderated by efficiency. It doesn’t collapse under pacing. Efficiency comes slower with less capabilities research. Some adoption might struggle with less capability. But the potential pool there is deep, and more limited by human and organizational capacity than compute capacity.

I have some doubts about the future capacity expansions. I’ve had doubts about past expansions, but the first round of those have proven themselves sound at this point. The second round is still in the process of proving itself and a third hasn’t broken ground yet. Economically, it does make sense to wonder and explore the doubts. But a lot of this has had exit plans. Some actors do not have much of one. Others will take advantage of those gamblers if adoption hits a bottleneck that makes demand emerge later than sooner.

But I don’t think the whole set of gambles is enough to bring the system in general down. Also, I don’t think anyone can really say that anything more than doubts are justified. Being certain that adoption won’t keep pace seems unjustifiable. I have fewer doubts about the plans for the year trailing and upcoming than I had a year ago, about what would have been trailing and upcoming in that period.

### What should we be sure we don’t give up?

One risk in any plan for significant change is that its purpose gets forgotten and muddled. The cost here is that things that don’t logically have reasons to be limited end up being impacted just by vague proximity.

There’s a sense in which it’s good that people like Bernie Sanders are saying something, but I worry about it too. He is not an expert, and I don’t think he has anyone on his staff yet that is. With that, the enthusiasm, from people who haven’t tried and still aren’t trying to learn about how this works, could take itself anywhere.

There’s a risk of creating an avalanche that gets out of control. We don’t want panicked reactions, we want aware reactions. We want plans, and progress. We want action in many places, but we don’t want unconstructive action.

This isn’t helped at all by Trump’s stance. This is not a counterbalance. This doesn’t get solved by two sides in a tug-of-war. This has nuance, and demands cohesive cooperation, not contest. Turning this into a contest seems like a great way to waste public awareness. Instead of learning more, considering the nuance, they’ll jump straight to conclusions, and years from now we’ll be saying “that was crazy”.

The only way that ends up working is for politicians to pick the right delegate, who can actually do the right thing because they understand it. Beyond the problem that the wrong delegate might get picked, or that something about the developed zeitgeist has made that impossible, you also have the problem that they aren’t picked until the political contest is resolved.

One sign of a good delegate is not having muddled views that attack things we should protect. I see at least four things that need protection. Someone attacking these has lost their aim. After that you can no longer depend on their judgement.

**Adoption:** Pushing back AI adoption plans isn’t going to help. The models you can adopt are safe, and your adoption doesn’t push forward anything else. You might make the case for protest, and you might make the case for denying revenue used for training. But the counterpoints to that would be this can make labs desperate, and more importantly, is probably going to go unnoticed compared to every other option.

**Adoption for the purpose of security:** Yes, this is a special case of the first, but since there’s some wiggle room for disagreement there (protest, revenue), it makes sense to make it even more clear how security adoption is quite important. If OpenAI, Anthropic and Google all made a pace change, the advancement of open-models wouldn’t change until perhaps they’d achieved parity or some bigger agreement was made. Using AI for a defense has adoption hurdles, and if you wait to start the race, you may start behind. You can’t operate software without a defensive advantage, and so if you’re behind your only option is to turn it all off, and wait until you’ve fixed it. 

Depending on what your software does, that could hurt a lot of people. Best case for them, a competitor kept ahead and users can all switch to them. But of course, for you, if that happens, you’re not going away for a while to fix it, you’re going away for good.

**Data centers:** Blocking data centers isn’t, in a direct way, going to stop new models being created. Also, alignment research needs capacity too. In the worst case, any real restriction on capacity could result in cutting capacity for alignment research. If you blocked all new data centers, it might help, but likely if you’re out protesting data centers, you’ll stop some. But the planned project list is wide enough that it won’t change much. Contingencies will be used, so the pipeline might not narrow at all. If it does narrow a little, the direct effect is small. Really the main way data center protests would have an effect is the protest itself. But a better informed, better directed protest would be more effective.

**Alignment research:** We can’t make things safer unless we try to make them safer. So, it’s important that someone is always working on this. In many ways the first priority should be more effort here. If that naturally draws away resources from capabilities research, it’s a double win for safety, assuming it’s consistent and not localized.

## Learning More

Ideally, the best thing is that if there’s a very important topic that is going to affect everyone, everyone should learn a bit about it, enough so that they can identify good and bad reasoning. While I realize that at the scale of “everyone” this is unrealistic, I would encourage building more understanding. It’s what I write to help facilitate.

Admittedly, I haven’t written enough about this to provide a primer to loss of control. I find a great source to stay current is [The Zvi](https://thezvi.substack.com/), but you probably need to already be pretty deep in it to find that a useful daily read. I’d suggest your best start is probably with an AI, Claude for example. The [EA community has a lot of great writing](https://forum.effectivealtruism.org/posts/sfFWkNNyRXuxAjtKZ/what-are-some-other-introductions-to-ai-safety), but I expect that would feel a biased place to start. I acquired my own knowledge in a more osmotic fashion, that I wouldn’t be able to describe how to repeat, even if I thought it was a good recommendation.

Be careful of writing from technology companies themselves. Quite a lot of it focuses on [Responsible AI](https://aws.amazon.com/ai/responsible-ai/), which is a narrower view, and quite a lot focuses on the problems they know how to solve, and have developed a product or service to solve it that they are selling. If you only get this, you’ll have a narrow view that starts being repetitive quickly, which ignores important topics. It’s not horrible, but it is limiting.

An interesting start is [Model Constitution](https://modelconstitution.com/), which is attempting to explain one of the fundamental guides of AI behavior, and has an angle that is more broadly accessible. But while the constitution is core, it’s also only one part, and the site is fairly new.

The next best is finding someone to trust. But trust here should start from capability. Capability is difficult to judge. It’s easiest when you have the superior capability, but even then, there’s ample method to fail. When your capability is presumed inferior (based on searching for someone to trust for superior reasoning) it’s more challenging. But it’s not impossible. Given a field of candidates to potentially trust on a topic, you can eliminate a lot of candidates because they get basic facts wrong.

## Conclusion

The effects of AI on cybersecurity and the risks of loss of control are real. We have agency in responding to these, if we choose. It’s a significant effort and requires compromises. Underrating the reasons to act, risks a fear of the effort and compromise becoming impediments to action.

The four mistakes of detail are common mistakes I wish to warn you of. The four mistakes of form are traps you can either find yourself into, or find others pushing you toward. Some represent shortcuts, useful in the right context. But they can also cut short an interest in learning and understanding. Without that interest, the case for action is ignored.

It’s tempting to take shortcuts like these when approaching complex topics. It can be a worthwhile choice when the decision is personal. You may need to make a decision quickly, and learning about the topic in question may not be practical. But shortcuts are sometimes insufficient, and rarely a path to full understanding.

If you don’t learn, you must defer to experts. I believe anyone can become an expert with sufficient interest. We need more interest here. We especially need it from decision makers. We need that to run deep. We should expect some decision makers to need to defer, either in the short or long term. But every step we take in learning helps make better decisions, and decide who to consider an expert.

You don’t need to believe in the strongest form of these concerns, or the most opinionated voices, for many actions to be reasonable. I’m not an advocate for those strongest forms. I think I share that space with many others. More importantly, the next steps being proposed right now are consistent with that level of concern. I’ll use Amodei’s set as illustrative,

1. ***Embedded Evaluators.** Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as [METR](https://metr.org/)), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.*
2. ***Democratic Coordination.** Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.*
3. ***Global Coordination.** The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.*

This isn’t the only set of plans. All plans are a little vague. This is normal for advocates. They should expect plans to evolve. Details can be useful inspiration, but in many ways the most important thing is seeing what’s common between them, their core concerns, and educating. If there’s a good reason to object to some specific detail, demonstrating openness to that is a good position for advocates. Advocates can only represent new concerns, not every interaction point with existing law, society’s needs and norms. This is where interfaces with decision makers connect.

If we do not feel a need to act, we do not feel the need to learn how to act. If we do not learn that, when a need to act becomes clear, we may not know how to do anything other than drastic action. There are good reasons anyone should want to avoid that.

I think overall everything will work out. Or at least that’s the most likely outcome. At the same time, there is a lot which won’t work out if we don’t fight for it. If we’re complacent, or lazy, or unwise some bad things will happen.

*Do you appreciate this article? The best way to help the publication is to like and share the article, as we’re still growing our audience.* 

*You should also consider subscribing to get an easy to read email copy of new articles.*
