cd /news/artificial-intelligence/iii-anthropic-reasoning-has-issues-w… · home topics artificial-intelligence article
[ARTICLE · art-85276] src=lesswrong.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

III. Anthropic reasoning has issues with infinite worlds; D-SIA can fix this

A new variant of the Self-Indication Assumption (SIA), called Distributional SIA (D-SIA), fixes most issues with infinite worlds in anthropic reasoning, according to a LessWrong post. Classical SIA (C-SIA) breaks down in infinite worlds due to divergent reweighting and non-updating, but D-SIA, with the right choice of prior, remains well-behaved in all cases. The post argues that D-SIA avoids problems such as infinite equal reweighting and non-updating in our universe, though it may still face divergent reweighting unless using modally controlled priors.

read23 min views1 publishedAug 3, 2026

SIA is true when there are no duplicates to worry about, and though it has issues around duplicate creation, so does every other theory of anthropic probability.

More seriously, though, it has serious problems with worlds with infinite numbers of agents. Ironically, SIA is often used to argue for infinite numbers of agents – so it breaks in the very worlds that it advocates for.

I’ll present D-SIA (Distributional SIA) which fixes most of these infinite world issues (with the right choice of prior, it fixes them all). However, it doesn’t argue for infinite worlds in the same way that classical SIA (C-SIA – classical or counting SIA) does.

SIA upweights each world by a factor of the number of agents (subjectively indistinguishable from the reasoner) that it contains. This leads to a series of problems in worlds with infinite numbers of agents (“infinite worlds”, colloquially):

For 1, define world as containing copies of the agent. Give a prior probability of for . This is a well-defined prior. But SIA will update world ’s weight to , which has no finite sum. So this cannot be renormalised to a probability distribution over these worlds. Note that though the number of agents is finite in all worlds , the expected number of agents is infinite; SIA breaks down in that case. Problem 2 follows from the fact that two countable infinities are the same size, so neither is larger than the other. For 3, the conditions imply that, whatever any agent observes, there will always be infinitely many copies of them in the universe, so 2 will always apply. Our universe being infinite and fundamentally random at the low level gives 4.

For 5, just compare with a world that is actually exact copies of world – maybe is an algorithm and runs that exact algorithm times. Or maybe contains causally disconnected identical galaxies. In both cases, SIA will provide an -to- boost from to , despite them being subjectively indistinguishable, with being arbitrarily large, potentially infinite. If there are non-observed facts that are different in the two worlds, their probabilities will also get similarly boosted. I’ll propose a new variant of SIA, called D-SIA (distributional SIA). It is different from traditional SIA, which I’ll rename C-SIA (counting SIA) when I need to distinguish the two.

D-SIA naturally solves the direct problems of infinite worlds – it’s fine with two worlds containing infinitely many agents, and can compare two infinite worlds with ease. So 2, Infinite equal reweighting, and 3, Infinite non-updating, are not a problem. It also avoids 4, Non-updating in our universe: it’s fine in our universe, once defined. Its priors may be chosen to not reweight on invisible differences [1] .

Strangely, it still can have a problem with 1, Divergent reweighting. It will find convergence in many situations where C-SIA doesn’t; but not all of them. But there is a specific class of priors for D-SIA – the -modally controlled priors – where D-SIA will always be well behaved, on every world and every comparison. And any set of priors can be approximated from below by -modally controlled priors by letting go to zero.

In standard Bayesian reasoning, you downweight each world by how surprised you are by the observation.

So, for instance, if you have two theories, one that your friend is in a coma, and one that they’re doing fine. Then you receive a voicemail from them. If your friend is doing fine, this isn’t surprising at all. If your friend is in a coma, this is astounding – not impossible, it could be an elaborate fraud, but still astounding. So the “friend in a coma” theory gets a large penalty for the observation being very surprising.

This is why you update world with , the probability of making observation in world . If that probability is low, it means you were very surprised by the observation in world , and so world pays a hefty price.

Consider the Sleeping Beauty problem: in the tails world, there are two “yous”, one in room 1, one in room 2. Classically, you could consider yourself to be either one. And then you discover that you are in room 1. Again, classically, this would be considered a surprise, and should reweight the tails world, down, by a factor of 1/2.

But SIA behaves as if that surprise doesn’t count. And this is not unreasonable – after all, a version of you observing they are in room 1 is certain in the heads world, so it might be a surprising observation for you to make (assuming that there being a separate you makes sense), but it’s not a surprising observation to be made in that world.

Now, refusing to update on observed information is very wrong (and non-martingale) under Bayesian updating. So instead:

So, if you could be one of agents in world , then finding out you were agent number has probability . SIA prepays this by reweighting the probability of world , up, by a factor of .

Notice a world with agents behaves as if they had an extra anthropic mass of to go around, multiplying their prior probability. This gives the first attempt at distributional SIA:

This “works”, in that it can be used to resolve most of the issues with SIA and infinity. Just choose bounded masses for every world, and everything becomes standard Bayesian. If is the prior for world , you can choose finite-but-unbounded masses if the sum of the is finite, and everything remains standard Bayesian.

Of course, this only works by completely giving up. If you get to decree the anthropic mass, then there is no anthropic problem. C-SIA is mass SIA where the is uniform over agents and the mass is the count of the agents; that count is, at the very least, a reasonable non-arbitrary candidate for defining the mass.

In our universe, the cosmological constant gives a maximal size to any causally connected structure in the universe. The Bekenstein bound gives a limit on how varied any agent contained within that space can be. The Landauer principle limits how many changes an agent can make to its storage (the cosmological constant forces a minimum temperature to the universe – Gibbons-Hawking radiation from the de Sitter horizon – which forces a minimum expenditure of work for any change).

So the number of distinct agents is limited by some (enormous) finite number . So there are only finitely many future agents you could be. If future information won’t distinguish between two copies of you – because they will always see the same future history – then we could consider that not pre-paying to distinguish between them. This suggests:

A first objection: what if two agents start identical but have a small chance of later diverging [2] ? Should they really be treated as the same agent?

They’re not treated as the same agent; they’re treated as having a certain probability of being the same agent. You give probability to , the divergence world, and probability to , the no divergence world, add an anthropic reweighting to only, by a factor of . The total anthropic mass looks like (which is if divergence is impossible, and the full if divergence is certain).

However, future history SIA still has issues. It only works because the number of histories is actually assumed to be finite. And though it solves Divergent reweighting, since the reweightings are bounded above by , and Infinite equal reweighting (it will compare the relative probability of one full future history to the number of different full future histories in that world), it doesn’t address Infinite non-updating if we are talking about “full future histories” as possible observations: as soon as a world contains one copy of each full future history, the theory has nothing more to say.

And it likely fails at Non-updating in our universe, not because of Boltzmann brains (who won’t have much of a future history), but because our universe is likely infinite and stochastic, so each full future history will exist somewhere in the universe.

So, now let’s assume that our anthropic distributions are non-uniform; on each world , we have a distribution over who the agent could be. Now different agents have different probabilities, so the agent expects to pay different amounts to discover who they are (conditional on who they actually are).

So the expected probability, if the agent was just about to discover who they were, is . This can be written as the inner product , or (referred to as the Simpson index).

We can then define:

First observation: is always finite, even if there are an infinity of agent types. So no world gets a bad update, even if it has infinitely many agents in it.

You can still get a bad distribution over worlds, though. If is uniform over agent types, then , so . Thus uniform reproduce C-SIA, and the divergent reweighting was a feature of C-SIA sensibly defined over finitely many agents in each world; this D-SIA shares that weakness.

Note that the , and the definition of agent-types they are defined over, are important to the problem we’re working on. Why, for instance, do we not think that we’re Boltzmann brains? Well, Boltzmann brains are rare and their density is very small across our universe – most observers aren’t Boltzmann brains. But note that “most observers” is phrased in terms of, e.g., “most observers in a given volume of space” (a property we can define using distance as a measuring stick) rather than “most total observers” (an undefined property, as the ratio of two infinite numbers is undefined). Thus the will encode the structural assumptions we are making about the universe, agent identity, and the problem we are working on.

Consider this problem:

Question 1: What D-SIA probability should you give to heads or tails?

In the finite setting, it’s on heads vs tails, because there are twice as many agents in tails. But in the infinite setting, the number of agents is exactly the same – infinity. So either C-SIA gives up entirely, or it says that the probabilities of and are equal to their priors, thus .

For D-SIA, we can let each world have a probability distribution over types of agents, which in this case is the room number. So is concentrated exclusively on room 1, while is split between room 1 and room 2. But before we plough ahead and compute with those distributions, note how we got them. There are infinite numbers of agents in both worlds, but what we are saying is “we can pair off the rooms within a single galaxy in a meaningful way”. That’s what allows us to define the values in . So, by using these distributions, we are refining the original question to:

Question 2: What SIA probability should you give to heads or tails within each galaxy?

Now Question 1 remains meaningless, but Question 2 is actually solvable in this setup. As computed above, the of a uniform distribution over elements is , so and ; just as in the finite case, we’ve reweighted the tails world, up, by a factor of two.

But now consider:

Question 2 becomes meaningless, because not all galaxies have the same proportion of heads or tails. Instead we want to aggregate the local information, averaging it out across the universe. And we’ve implicitly set that up already. Because in the definition, there is the term “one in a trillion galaxies”. What is defining what “one in a trillion galaxies” actually means? Reasonably, we might be thinking of average density: if we picked any point and expanded a ball around it, the proportion of erroneous galaxies would tend to one in a trillion.

So we get question 3:

Question 3: What mean SIA probability should you give to heads or tails within each galaxy, given a density-based way of aggregating the galaxies?

This is a local, aggregated probability question. It concerns local properties (Sleeping Beauty room numbers are defined within each galaxy) aggregated across the world in a spatial-density way that makes sense [4] .

For Infinite Sleeping Beauty with errors, the distributions aren’t hard to define; picking a galaxy uniformly by density then picking the agents: with everything on room 1, as before, and with on room 1, and on room 2. So is still while is about . So the final probabilities are about and . In both variants of Infinite Sleeping Beauty, the distributions had two types: in room 1 or in room 2. So there are two agent types – two sets, each containing individual agents.

But what if we cared about the differences between the many, many agents here? What if we, for instance, put a probability distribution across the infinitely many galaxies? We’re taking the previous two categories of agents and splitting each of them into infinitely many subcategories – a refinement.

Then there is an invariance result with even refinements. Pick any distribution , and split every agent type according to , and then D-SIA is unmodified by the change. This is because the product distribution has . So each world picks up a constant factor of reweighting ; upon renormalisation, this vanishes.

In general, the agent type is defined to be sufficient for the problem at hand – thus the two agent types in the infinite Sleeping Beauty problem. But in general, for updating on observed information, we don’t want the agent types to split on observations.

Recall that Future history SIA used the agent’s possible full future histories as agent types, while assuming that there were a finite number of each. We can take that definition, while dropping the finiteness requirement.

This is the coarsest definition of agent types that will never split on any observations, so is always a potential candidate for defining agent types. It also gives a cross-world identification of agents (which D-SIA doesn’t give or need): two agents in different worlds are the same if their full histories are the same.

This agent type definition solves the Reweighting on invisible differences issue: the avoid such reweighting iff their agent types are (a coarsening of) the full future history agent types, or are an even refinement of such a .

So far we haven’t been considering agent actions, but it’s now relevant. Suppose that we had four currently identical agents, who are defined by their coordinates: left versus right, top versus bottom. They have to choose either vertical or horizontal. If they do, they find out which coordinate they are, but only for that one coordinate. So if they choose vertical, the two tops will see the same thing and the two bottoms will too – but left versus right is never split. Opposite if they choose horizontal.

Then we can reasonably consider them four agents; the formalism is that a unique agent is identified not with a unique full future history, but with a unique map from deterministic policies to full future histories. Each of these agents has a different such map.

So the genuine coarsest agent types that always avoid Reweighting on invisible differences are grouping them by maps from deterministic policies to full future histories.

So the D-SIA reweighting is always finite, solving most of the problems of C-SIA. Updating is standard Bayesian on the re-weighted prior. In terms of a deterministic world , with agent prior and weighted prior , updating on observation sends to

[5]
, and sends to (which we’ll write as ).

Thus D-SIA avoids Infinite equal reweighting and Infinite non-updating. All that remains is Divergent reweighting. But before that, let’s check D-SIA on Boltzmann brains.

Boltzmann brains are extremely rare (there was probably not a single one in the whole history of the observable universe). Let be the total density of all Boltzmann brains in the universe, and assume that there is at least one agent type considerably more likely (in density) than a Boltzmann brain – so (in practice, there are many such agents). Then even if all the Boltzmann brains were in a single agent type, they would contribute to , much less than the contribution from . If the Boltzmann brains are distinct types, their contribution would plunge even more.

Thus adding a small probability of relatively unlikely agent types to a distribution will make only a small difference to . Hence D-SIA doesn’t collapse in the presence of Boltzmann brains; it doesn’t have the Non-updating in our universe issue.

How can we aim to avoid divergent reweightings? There are two direct methods: restrict to finite agent types, or use modally controlled priors. That second approach can be used to approximate any priors from below.

We can use the full future history list of agent types, and import the assumption from Future history SIA that this collection of agent types is finite. Then if the number of agent types is less than , and hence all the reweightings are bounded by the same factor : the reweighting must define a proper probability distribution.

Now D-SIA works beautifully and simply. Note that the C-SIA version of full future histories would fail because it counted, and as soon as there was a single representative of the full future history in the world, it was the same as if it had ten trillion. So if all full future histories existed somewhere in an infinite world, C-SIA would fail to distinguish that world from another one with all the histories.

But D-SIA, using distributions rather than counts, has no such problems. Any other setup that restricts to finite types will also work with D-SIA.

The prior is -modally controlled if there is a such that, for all , .

Then, automatically, we have , and hence ; so all the reweightings are bounded by the same bound, and the reweighted distribution is automatically well defined (notice that the doesn’t have to be the same in different worlds – the condition is just saying that every world has at least one that has probability at least ).

Ok, but what if we don’t have a modally controlled ? Then we can replace any specific with , getting modally controlled . To do this, find a modal element for each world ; so for any other (as before, there is no reason to expect or require that will be the same in each world).

Then if , . Otherwise, rescale to be modally controlled by : let and for all other , where . So an extra has gone to , to boost it; the remaining elements are then scaled down to compensate.

Thus we can replace any with , and the new prior will always have well-defined D-SIA. Then we can let shrink towards (with D-SIA well defined at every stage) to approximate .

There are some choices involved in defining D-SIA. The full future history agent types seems a nice definition of agent types, but it might seem clunky and overly subjective. The themselves are often defined in terms of the problem being addressed; non-even refinements of lead to different results (in fact, a duplication event can be seen as a non-even refinement of in the middle of an agent’s history; the inert duplication variant can be seen as replacing that non-even refinement with an even refinement, which doesn’t change the anthropic probabilities).

Let’s look at what happens if an anthropic agent uses D-SIA to reweight their prior and then observes , versus waiting to observe and then reweighting their prior. Say that the moment when the D-SIA agent reweights their prior is the anchor moment. So we’re comparing D-SIA at different anchors.

Let’s follow a specific world through both anchors; relevant prior information is and . We’ll ignore normalisations of reweightings here, since, as long as normalisation always remains possible, it doesn’t matter if the normalisation happens at each step or only at the end.

First, goes to ; then this is further updated to . The gets updated to (otherwise known as the conditional distribution ). So we have .

Now let’s apply the observation update first. This changes to , removing all agent types that didn’t observe (and removing all where no agent type observed ). Upon doing the D-SIA reweighting, this goes to .

So are these two quantities equal – is the same as ? For C-SIA and as uniform priors, they certainly are: assume agent types in observed while didn’t, then , and .

But they aren’t equal for generic, non-uniform . However, if is their ratio , then has value in expectation [6] .

Since is not identically , different anchors have different distributions: waiting for an extra observation shifts the probabilities. But since is in expectation, this shift is martingale: delaying the anchor doesn't change the expected weight on . And note that could have been a single observation or a whole sequence, and the argument would have been the same – observations still partition the agents in deterministic worlds – so delaying for any length of time doesn’t change things.

Now is identically for C-SIA for all observations, which means that, when considering C-SIA, people didn’t ever need to think of quantities like , which may have made finding D-SIA harder and explained why it wasn’t discovered earlier.

Needing an anchor is a weakness of D-SIA (even with the martingale result), since it needs an extra choice. There’s always the initial history anchor available as a natural choice, however: push all agent histories backwards to their initial points, anchor there, and push the distribution to the present via Bayes. Alternatively, one could push to the future, to the end of the completed agent histories, and anchor the D-SIA there [7] .

So, we have D-SIA. And with the modally controlled priors, we can now compare the probability of any agent in any world. Finally, can we answer the deep questions of anthropics – is the universe infinite? Is there a multiverse?

And the exciting answer is... D-SIA is pretty agnostic about infinite universes and multiverses. It depends on the choice of priors. The agent types of the encode our judgments about diversity type – what counts as different agents, fundamentally? Do the identical agents of a finite universe, infinitely repeated, count as one agent or infinitely many? If you think of them as the same, don’t have a prior that distinguishes them; if you think of them as different, do distinguish them.

This is a case where it’s very clear the prior is doing the work. You can choose to consider these multiples as one, as distinct agents, or as some discounted sum. That choice, combined with an anthropic probability method that scales with multiple agents, is entirely determining the weights of these repeating universes.

If the universe were finite, we might hope that, maybe, we could “exhaust” all the observations in the universe and make the prior less relevant. But in an infinite universe, this can’t happen. In a universe where we have only observed a potentially infinitesimal fraction of its features, the prior is doing the work of translating those observations into probabilities. That doesn’t mean that we shouldn’t speculate about infinite universes. I think there’s a high likelihood that our universe is infinite, given General Relativity with an observed cosmological constant. The maths and the elegance of the formulas have been quite convincing to me. But my prior over elegant maths being right is probably higher than most people’s.

The same goes for multiverses, except in some cases, it’s even more obvious that the prior is doing the work. If we have a theory of multiverses of disconnected universes, then the following are completely indistinguishable: a) us being in that multiverse, in a universe that looks a lot like our own and b) us being in a solitary universe that looks a lot like our own. Anthropics and D-SIA can provide extra updates, but we have to have chosen to allow these updates in the first place.

Again, “prior-dependent” doesn’t mean “completely uninformative”. I can see arguments for using very fine and counting re-run copy universes separately; I can also see the arguments for coarser (such as the possible observed histories) and counting them as one. From a mess of priors and impressions, I feel a multiverse is probably likely. But arguments (and argument-relevant evidence) will probably change my views in the future. But I don’t expect evidence to ever do so on its own, because by definition, we’ll never see the other universes. So evidence has to be filtered through a prior or a theory in order for it to change my probabilities.

There’s a common argument, which is that SIA is clearly correct in finite cases, it scales by the number of identical agents in the world, so any world with a finite number of copies of you is dominated by larger ones, hence you must live in a universe with infinitely many copies of you.

Notice what the argument does: it uses C-SIA, a theory that breaks in infinite worlds, to argue that the world must be infinite. That’s a reductio against C-SIA, not an argument for infinite worlds.

As I said in my previous post, reasoning with priors and updates is extra difficult in anthropic problems. That’s because they tend to produce exceptionally large (sometimes infinite) updates towards one position or another. With that magnitude of update, the headline number is of little use; all the real uncertainty has gone into whether the model and the assumptions are correct. How certain are people of their prior and arguments? Maybe quite certain. How certain are they that they haven’t overlooked an alternative model with a very different conclusion? If someone is certain of that, then they are certainly wrong.

Priors may even stop being sufficient because the question of reality becomes unclear. In a certain sense, every object in a Tegmark level four multiverse exists. In a certain sense. How would you even begin to put a prior over that?

Its prior may also be chosen to deliberately update on invisible differences. ↩︎

Full non-indexical conditioning (FNC) exploits the fact that real agents in the real world will always start diverging. It is built on a nice insight, but is non-martingale – essentially the anthropic mass increases as the probability of divergence increases over time. If you applied it to future histories – to the final, as-diverged-as-possible agents – then it would be martingale and would be future history SIA. ↩︎

We can model a stochastic world as a probability distribution over deterministic worlds, so we can always pass from a prior over stochastic worlds to one over deterministic worlds. It’s not normally that useful to do so; but here, we need to carefully distinguish between updates to agent-position within a world and updates to the world probabilities themselves, and deterministic worlds make that much easier. ↩︎

But if a better way of aggregation is present in another universe design, then that method should be used. ↩︎

Here, is just the total probability assigns to the agent types in that observe . This is well defined, as being deterministic ensures that each agent observes or doesn’t observe , while the agent types definitionally do not split on observations. ↩︎

For each observation , write . Because the world is deterministic and agent types don't split on observations, each type observes exactly one , so the blocks partition the types and .

Now, for any that observes . So, since is summed over all the agents who make observation , this must be . So the ratio of early-anchor to late anchor weight is . Taking the expectation of involves multiplying this by , giving , and then summing over , giving . ↩︎

Why not both? Push to the past, to get the maximal number of agent types, then push to the future to get the final probabilities. ↩︎

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @lesswrong 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/iii-anthropic-reason…] indexed:0 read:23min 2026-08-03 ·