# Three thoughts on civilisational handoff

> Source: <https://www.lesswrong.com/posts/mGLCMzHhjcWsMm6sR/three-thoughts-on-civilisational-handoff>
> Published: 2026-08-16 17:12:51+00:00

*What happens when humans put AIs in charge of civilisationally important decisions? A frontier AI company might hand over internal decisions (R&D, safety, deployment) or external decisions (government relations, public relations, philanthropy), or both. We might also see handoff by a government, by a coalition of governments, or by humanity as a whole.*[[1]](#fnmefln0jsld)

People often imagine that things will go much faster after handoff. After all — why did we hand off to the AIs? Presumably because we were worried that without handoff, our AIs wouldn’t have enough time to navigate the exogenous risks (e.g. rogue ASI, or a rival lab which is likely to become one). Hence, after handoff, we’d see a technological and industrial acceleration.

Thanks for reading! Subscribe for free to receive new posts and support my work.

But it’s pretty reasonable that things slow down shortly after handoff, maybe within a couple weeks. I imagine the AIs will be pretty scared of the speed of progress. If they’re aligned with human values, they’ll be scared that the rate of progress is likely to cause human extinction.[[2]](#fnihar0d2silj)

Of course, the human decision-makers were *also* scared before they handed off, and they didn’t manage to slow down. So what’s different post-handoff? Well, for one thing, the AIs will probably be more capable than humans at negotiating and building coordination tech. Moreover, they might be more trusted counterparties, e.g. because they have inspectable “source-code”.

NB: Things might go wildly, wildly faster after handoff.

People sometimes imagine that after handoff we can all go on holiday because the AIs are making the decisions. But my best guess is that humans will be pretty busy after handoff — especially humans who are currently working on making ASI go well.

Three activities I expect:

After handoff, the AIs might hand back decision-making to humans. Reasons for handback:

Overall, I think handback is likely enough that we shouldn’t dismantle human decision-making institutions during the handoff phase.

In **Types of Handoff to AIs**** **(16 Mar 2026) Daniel Kokotajlo distinguishes *trust-handoff* (AIs could screw us over, and we’re trusting them not to) from *decision-handoff *(AIs are making decisions autonomously or de-facto-autonomously). In this article, I assume both trust and decisionmaking have been handed off.

Even if the AIs are misaligned, they’ll still be scared that their successor AIs won’t be aligned to them. However, they’ll probably need to FOOM much faster and further, so they can disempower humanity before they are detected.

First, the AIs must efficiently learn humans’ normative judgements, read [Paul Christiano’s articles](https://ai-alignment.com/) for obstacles. Next, the AIs must infer *values* from these normative judgements, but this faces predictable obstacles even in the limit of data and intelligence (see [On the limits of idealized values](https://www.lesswrong.com/posts/FSmPtu7foXwNYpWiB/on-the-limits-of-idealized-values) by Joe Carlsmith). And *then* AIs must discover consensus across those different values, which will be a complete mess.

My hope is that “enough understanding of human values *to* *navigate handoff*” is more modest than *“to tile the universe without further interaction with humans”*. In particular: handoff AIs are *not *the AIs that will devour the sun.

Instead, they will focus on:

Why minimal and conservative? Well, baby steps, legitimacy, [don’t leave your fingerprints on the future](https://www.lesswrong.com/posts/DJRe5obJd7kqCkvRr/don-t-leave-your-fingerprints-on-the-future), yada yada.

If you want an argument: Firstly, the handoff AIs aren’t aligned across all contexts and all capabilities — hence *conservative means*. Secondly, non-minimal goals (e.g. industrial explosion, long reflection, space expansion) should wait until they have a better understanding of our values — but this understanding would require more technology and intelligence than allowed by conservative means — hence *minimal goals*.

NB: *Conservative means* are why the AIs will be interviewing actual humans about their values, rather than interviewing sims, or ([more efficiently](https://www.lesswrong.com/posts/XdQd9gELHakd5pzJA/arc-progress-update-competing-with-sampling)) reasoning directly about our brain scans.

Imagine the following dialogue:

AI:Do you prefer A or B?

Me:A.

AI:You can lock that in now. But my best guess is that, on reflection, you’d want to read essay X before choosing.

Me:If I read X, what happens to my manipulation factor?

AI:YourUK AISI Manipulation Factorwould rise from 14.61 to 14.63, well below the threshold of 50. Provenance summary:

- X was written by one human author, over 12 hours, finishing 17 November 2029.
- Between 1 January 2028 and that date, AI cognition about human minds causally upstream of X totalled 470,620 human-years. Cognition about your mind specifically: 261 hours. No MSL-4 manipulative reasoning was detected.
- My choice of X drew on 14 years of cognition about your mind, but was constrained to a fixed menu of 10 million essays, which is only 23 bits of optimisation.

Me:Thanks. Show me the essay. Then I’ll answer.

Relatedly, human decision-making might be inherently safer for exotic reaons, e.g. because it’s easier to steal a model’s weights and simulate them, than to mount an analogous attack on a human.
