cd /news/artificial-intelligence/taboo-equilibrium-less-confused-fram… · home topics artificial-intelligence article
[ARTICLE · art-82504] src=lesswrong.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Taboo “equilibrium”: Less confused frames for research on AI bargaining

A LessWrong post argues that common frames for understanding AI bargaining problems are confused and proposes mechanistic explanations for why powerful AIs might fail to coordinate, emphasizing the need for research on safe Pareto improvements (SPIs) to mitigate AI conflict. The author, identified only as a LessWrong contributor, recommends reading prior posts on AI bargaining and game theory, and breaks down ways agents can guarantee coordinated demands, noting that conditioning on the other's demands can lead to infinite regress.

read13 min views1 publishedJul 31, 2026

To understand why powerful AIs might get into conflict, and ways to mitigate it, we need to understand bargaining problems: situations where multiple agents have different preferences over

I’ve come to suspect that certain common frames on bargaining problems are confused. Here, I’ll explain why, and which frames I think are better. One motivation for this is to hopefully help others make progress in research on safe Pareto improvements (SPIs), which are among the most promising approaches to mitigating the downsides of AI conflict, in my view.

In this post I’ll:

I recommend reading the following beforehand (at least the linked sections), but it’s not strictly required, since the essential parts for our purposes are in the quotes below.

“A high-level model of AI bargaining”, which spells out what kinds of strategic situations we’re aiming to understand:

“Responses to apparent rationalist confusions about game / decision theory”****:

If we want to make reasonable (even if coarse-grained) predictions about how AI bargaining will work, we should aim to **mechanistically explain ** bargaining behavior. This may sound trite. But in my experience, it’s easy to neglect mechanistic explanations and fall back on models or heuristics that don’t fit this context.

To see what I mean, we’ll start with the central question about AI bargaining:

We’re supposing that our AI bargainers Alice and Bob are technologically advanced enough to make fully credible, conditional commitments (via programs). And they can

[overcome private information problems]. Then, why wouldn’t they always coordinate on compatible demands?

I’ve partly answered this with the “ Ex post optimal =/= ex ante optimal” quote above.

Instead, we should get clear on what mechanistically explains miscoordination — because these explanations will inform our predictions about other aspects of AI bargaining, e.g., when agents will individually prefer to use safe Pareto improvements (SPIs). So let’s break down different ways we might think Alice and Bob can guarantee their demands are coordinated, and how each of them can fail. (This is a weaker claim than “miscoordination of demands is inevitable or rational” or “risking costly conflict is inevitable or rational”, both of which seem false.)

One way for Bob to make conflict less likely is to condition his program choice on Alice’s demands, that is: before choosing his program, first learn which demands Alice’s program is likely to make against different programs (e.g., by accessing a trusted server where Alice’s program is registered), then update on this information.

[3]
If Bob does this, he’s more likely to make demands compatible with Alice’s.

But suppose Alice’s program can make different demands depending on whether Bob accessed the server. Then:

(This is the dynamic captured by the “partial answer” above.)

Bob now wonders which program he should blindly choose. He considers a program of this general form: *“Make my initial demands with some ‘first-order’ program. Check if the demands made by the first-order program and Alice’s program against each other would be incompatible. If so, make lower demands than the first-order program’s demands.”

  • We’ll say a program of this form But Bob faces the same problem as before! Alice has an incentive to give her program the clause: “If Bob’s program conditions its demands on incompatibility, then demand more than otherwise.” So if he’s willing to accept the risk, he might commit to a program that doesn’t condition its demands on incompatibility.

Finally, even if Options 1 and 2 fail, perhaps Alice’s and Bob’s demands can be acausally coordinated? The idea is: Suppose Alice and Bob follow acausal decision theories. Then, they each might regard their choice as determining the other’s even without communicating, if they think they make decisions according to sufficiently similar procedures.

For our purposes, let’s say an agent’s decision procedure is the process that generates their choice of program given their evidence. To quote an objection addressed by the “Rationalist confusions” post: Virtually all agents who are sufficiently capable to enter high-stakes bargaining interactions will converge on the same decision procedure, and reason that their decision to [commit to a fair demand]

[logically causes]their counterparts to do likewise [or is strong[evidence]that they’ll do likewise]. In that same section of the post, I explain why this, too, isn’t guaranteed. This quote captures most of the problem (footnote added):

If the reason to expect convergence of decision procedures in bargaining problems is that we expect selection for the same decision theory, then …merely sharing the same decision theory that you consult as an ideal doesn’t mean you’ll share a whole decision procedure(with respect to the given problem). For example, you might not share how you model the decision problem, what kinds of evidence about other agents are most salient to you, how you approximate ideal Bayesianism, etc.[[6]]

So, we’ve got a mechanistic story for why AI bargainers might not coordinate on compatible demands: Bob anticipates that Alice has an incentive to use a program that bargains harder if he tries to avoid incompatible demands — either via conditioning his program choice on her demands, or conditioning his demands on incompatibility. (And vice versa.) And there aren’t sufficiently strong selection pressures toward very similar decision procedures.

But what does all this have to do with how we should think about AI bargaining research, more broadly? Here are two general implications.

Standard game theory, including lots of bargaining theory, predicts that rational agents will play a Nash equilibrium, where each agent’s strategy is optimal against the others’ actual strategies.

But in the kinds of bargaining problems we’re focusing on here, this is a very strong condition on its face! There’s a wide range of possible conditional commitments the AIs might implement. So realistically, AIs will have nontrivial uncertainty about each other’s strategies.

And here’s the payoff of the previous section: There’s an incentive not to resolve this uncertainty before locking in one’s program. That is, if an agent gains enough information about their counterpart for the prediction of equilibrium to make sense, this could leave them in a worse equilibrium (for themselves). I’m not aware of any standard justifications for predicting Nash equilibrium that explicitly acknowledge this kind of incentive. This should make us suspicious of such predictions in the AI bargaining context specifically, even if we’re inclined to defer to the equilibrium-first framework in general because it’s very well-established. (For what it’s worth, skepticism about equilibrium has academic precedent in epistemic game theory (h/t Jesse Clifton).

[[7]](https://www.lesswrong.com/feed.xml#fn-C2bMhHCctmHJH4zzA-7) [8]
)

Now, “nontrivial uncertainty” alone doesn’t imply the AIs won’t be in equilibrium. Shouldn’t we still expect advanced AIs to have beliefs about each other that are reasonable enough to avoid miscoordination? [9] I’m not sure we should. Here are some defenses of that view and why I don’t find them convincing:

So I remain skeptical that we should expect AI bargainers to play Nash equilibria. But we might ask, what else is there?

I think the way forward is to stop looking for game-theoretic “solution concepts” per se. Instead, we should model the agents as individuals making decisions under uncertainty, and see which strategic dynamics follow from plausible properties of such decision-making. Here’s an eloquent statement of this frame:

There is no special concept of rationality for decision making in a situation where the outcomes depend on the actions of more than one agent. The acts of other agents are, like chance events, natural disasters and acts of God, just facts about an uncertain world that agents have beliefs and degrees of belief about. The utilities of other agents are relevant to an agent only as information that, together with beliefs about the rationality of those agents, helps to predict their actions. (

[Stalnaker 1996]: 136) That is, if we’re going to make predictions about agents’ strategic decisions, I suggest we look at what kinds of beliefs they might have about each other, and why. This is hard, of course. But retreating to Nash because of theoretical tractability would be looking under the streetlight. And I think we can at least say some coarse-grained things about what kinds of beliefs are plausible. Examples of this methodology:

Recall the arguments in the “Mechanistic explanations of miscoordination” section. How did these arguments work? Well, we considered some actions Bob might take, asked what Bob might believe about how Alice’s program would likely respond to each action (and why he’d believe this), and checked which action Bob would expect to do best given these beliefs. By “expect to do best”, I mean to include various forms of weighing up actions’ possible consequences by their value and likelihood, not just literal expected utility maximization.

That’s it. We didn’t say anything like:

These sorts of claims are heuristics, not mechanistic explanations — at least, not without saying more about why we should expect Bob to reason in these ways. [10] They don’t tell us precisely why Bob might not (say)

In principle, heuristics can capture considerations our current mechanistic models leave out. But we at least need to point to reasons to trust that a given heuristic really is tracking such considerations.

And obviously, sometimes imprecise shorthand is convenient. For instance, I find the concept of “bargaining power” helpful for quickly conveying some intuitions here. But we’ll get confused if we take the shorthand’s connotations too literally when working out its implications — e.g., predicting that any opportunity to mitigate conflict comes at the expense of “bargaining power” (SPIs are a counterexample!). Disambiguating heuristics about bargaining helps us avoid talking past each other, and shows us when the implications we’re drawing from heuristics don’t actually seem to follow from (good) mechanistic explanations.

As an example, [11] imagine Alice commits to participate in some SPI with Bob if his program meets certain conditions (the details of which don’t matter here). Bob doesn’t know this yet, since he has stayed blind to all of Alice’s strategic decisions thus far. Now suppose we explain bargaining problems by saying that Bob “reasons as if he moves first”. We might end up concluding: “Bob might want to commit to a program that ignores all offers to participate in SPIs. After all, if he conditioned his program’s output on such an offer, he’d be moving second, which is risky/exploitable

I don’t find such an argument compelling on its own. It’s possible that Bob expects to do worse purely by conditioning on information about Alice’s program, no matter what specific kind of information it is. Yet it’s not clear why exactly he’d be likely to expect this. Talking about “moving second” doesn’t tell us that.

Some examples of analyzing expected consequences in multi-agent interactions, without relying on unscrutinized heuristics:

Thanks to John Massey, Tristan Cook, Lukas Finnveden, Cadence James, Matt Hampton, and Paul Knott for helpful comments.

Note that “mechanistically” doesn’t have anything to do with mechanism design. ↩︎ This dynamic might remind you of the “commitment races” problem. I think the commitment races frame can be helpful in some contexts. But it’s also loaded with associations that my argument here isn’t committed to (pun not intended), and I’m not aware of public writings precisely spelling out the argument in the rest of this section. ↩︎

Or, if he follows something like updateless decision theory, then he’s deciding whether to act as if updating on this information. ↩︎

Commentary on some other possibilities here: (i) Alice might also be capable of conditioning her program’s demands directly on whether Bob conditions his program choice on her demands, rather than on the crude proxy “did Bob look first?”. If so, this would presumably be a more robust option than using the crude proxy. (ii) Suppose Alice can’t design a program that conditions its demands on whether Bob looked first at all. Here’s another thing she could do with the same strategic consequences: monitor for whether, and when, Bob accesses the server — once he has done so, lock in a program that always makes high demands. ↩︎

If you’re familiar with “renegotiation programs”, this might look familiar. But renegotiation programs are importantly different. First, a renegotiation program rn(p) doesn’t adjust to the counterpart unconditionally: It only behaves differently from p if, against p, the counterpart’s program would’ve made the same demands as its own first-order program would. So the failure mode in Option 2 doesn’t apply to renegotiation programs, because a counterpart who demands more against rn(p) than against p thereby disqualifies themselves from the renegotiation. Second, in the case where rn(p) “behaves differently” from p, it doesn’t necessarily fall back to lower demands in particular. ↩︎ Something I left out of that post: Even if the agents share all these things, they might also not have access to the same evidence, when they decide what programs to commit to. ↩︎

Bicchieri (1995) (emphasis mine): “Yet even when common knowledge is present, there are games in which this much knowledge is not sufficient to infer a correct prediction of the opponents’ choices. Since playing a Nash equilibrium involves correct expectations on the part of the players, one might impose the additional epistemic requirement that the beliefs that players hold about the strategies of others are common knowledge. I shall argue that common knowledge of beliefs is an implausible assumption. Moreover, the assumption does not guarantee that players will end up with correct expectations if they have mutually inconsistent beliefs to begin with.” ↩︎

Pearce (1984) (emphasis mine): “The most sweeping (and, perhaps, historically the most frequently invoked) case for Nash equilibrium theory in such circumstances asserts that a player’s strategy must be a best response to those selected by other players, because he can deduce what those strategies are. Player i can figure out j’s strategic choice by merely imagining himself in j’s position. But this takes for granted that there is a unique rational choice for j to make; this uniqueness is not derived from fundamental rationality postulates, but is simply assumed. Furthermore, any argument suggesting that player rationality, combined with the structural characteristics of a game, inevitably renders all but one outcome ‘impossible,’ leads to conclusions that contradict widely accepted notions of ‘perfection’ (Pearce [19]). Once one admits the possibility that a player may have several strategies that he could reasonably use, expectations may be mismatched. Player i’s strategy will then be a best response to his (possibly incorrect) conjecture about others’ strategies, not the actual strategies employed. … [W]e are interested in analyzing many situations for which no precedents exist (such as nuclear wars between superpowers) or in which continual changes in relevant variables (technological breakthroughs, new legislation, and so on) preclude prediction based on tradition. It then becomes crucial to understand precisely what are the implications of players’ information and rationality.” ↩︎

Separately, if you’re familiar with Bayesian Nash equilibrium, you might be wondering why we don’t just use that instead. That is, we could model the agents as (a) uncertain of each other’s private information, but (b) still best-responding to each other’s

For example, if humans and current AIs find some crude strategic heuristics intuitive, that might be evidence that future AI bargainers will, too. I find it hard to say how strong this evidence is. But I think we should be explicit about when we’re making this kind of inference, and not conflate it with an argument from “rational” incentives alone. ↩︎ I made this up for illustration; it’s not from any actual researcher as far as I know, though it’s somewhat related to Soto’s discussion here. The real examples take more context and unpacking than makes sense for a post of this scope. ↩︎

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @lesswrong 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/taboo-equilibrium-le…] indexed:0 read:13min 2026-07-31 ·