Sleeping Beauty as a Mind Killer

Sleeping Beauty (SB) is a very popular logical puzzle, and there is an enormous volume of writing on the topic. No one can read it all.
Here I suggest that the SB problem was naturally selected to become maximally philosophically inflammatory. As a result, it loses much of its explanatory potential. If a correct answer exists, it is buried in tons of literature and depends on a number of assumptions.
The science-fictional setup does not help either. There are no practical situations in which powerful amnesia is used without damaging reasoning abilities. There is an analogue of SB involving twin brothers, but it has important differences: no sequentially appearing tests, such as Tuesday–Tails following Monday–Tails, are possible.
There are 161 posts about SB on LessWrong alone, compared with 3,200 about superintelligence.
The best minds, many of whom also work on AI safety, are spending their time on a puzzle that future generations, if any appear, may see as analogous to counting angels on the head of a pin. That problem also has depth: it requires calculating the smallest invisible thing, a task that could not be solved without a theory of light at the time.
SB simultaneously tests several ideas:
1. Probability vs. credence concerning a given toss.
2. Path-dependent identity across Monday and Tuesday under Tails vs. state-dependent identity across the two Mondays.
3. Different ways of aggregating bets.
4. Actual copies vs. possible copies: days vs. coin outcomes.
5. Whether possible copies should count as real in some sense.
6. Changes in the reference class after a new question, as in Bostrom’s hybrid model.
7. Prior policy vs. individual action.
8. A one-shot game vs. repeated games.
9. First-person vs. third-person perspectives.
10. Amnesia vs. copying.
11. Many-worlds interpretation (MWI) vs. classical models.
12. Different decision theories.
13. Bayesian vs. frequentist probability.
14. The nature of self-locating beliefs: SSA vs. SIA.
The perceived simplicity of SB conceals substantial complexity. It tricks the mind into proposing solutions that address only one of the distinctions above.
Below are several hidden caveats, or complexity bombs, within SB.
1. Probability realism is problematic
SB assumes “probability realism”: probabilities are real things, and we can correctly infer them.
Probability can typically be tested through frequencies or betting. However, SB is constructed so that both frequency-based and betting-based approaches can distort the result.
The frequentist approach does not work straightforwardly because, if the SB experiment runs only once, its outcomes are mutually exclusive with respect to the coin toss, though not with respect to the day. Under Tails, there are no instances of Heads. This supports the halfer position. If we run SB many times, however, the observations are no longer mutually exclusive and converge toward the thirder position. If the experiment runs only twice, the result is more complicated; see [Bostrom’s hybrid approach].
Beauty’s optimal betting behavior can also be calculated, but only if we assume that the experiment occurs many times. We also need assumptions about counterfactuals and about who receives which benefits.
Double-Extreme Sleeping Beauty
It is possible to imagine a one-shot SB experiment in which the stakes are extremely high:
God creates, just once, an unfair coin with a 0.999 probability of Heads.
If Heads, He creates one copy of me.
If Tails, He creates one million copies of me.
I have one attempt to guess the result of the coin toss. If I am wrong, I will be tortured and killed. MWI is false, and no one will ever repeat the experiment.
Tails or Heads?
When the stakes are this high, it becomes much harder to feel certain about what to do. In the ordinary problem, the difference between 1⁄3 and 1⁄2 may seem negligible. In Double-Extreme Sleeping Beauty, an error means almost certain pain and death.
SIA and the problem of exact copies
SIA and expected-utility reasoning favor Tails. The difficulty in applying SIA here is that the copies under Tails must be different, so that we can count them as distinct objects drawn from a pool of possible objects.
This can be illustrated by simplifying SIA into a box-of-souls problem. Imagine a box containing possible souls, each numbered from 1 to 100. I toss a coin. If Heads, I take one soul from the box; if Tails, I take ten souls. In this setup, soul 17 is ten times more likely to be selected under Tails, so any selected soul can treat the fact that it was chosen as evidence for Tails.
However, if the souls have no numbers, or all have the same number, 17, then soul 17 gains no information merely from having been chosen.
Thus, SIA is not simply an assumption but a provable theory that works only for different souls. In SB, however, Monday–Tails and Tuesday–Tails are internally indistinguishable. Perhaps we could simulate SIA in SB by assigning both states different random names before awakening. There is also a view that subjectively indistinguishable souls should still count as different: haecceitism.
In the “normal” SB problem, the awakenings are exact copies from the inside, and we do not know whether the universe is finite. Therefore, SIA may not be applicable to SB. Moreover, although SIA may be provably valid under certain assumptions, it can become self-defeating: it immediately favors an infinite universe in which all possible observers exist, after which we must return to SSA. See [SIA Becomes SSA in the Multiverse].
2. A coin’s intrinsic probability vs. credence about a particular toss
There is a subtle difference between the probability that a coin will land Heads, which is a property of the coin, and our credence that a particular past toss resulted in Heads.
We know that the coin has a 1⁄2 chance of landing Heads. Before observing the result of a single toss, we assign P(H)=1/2. We may then collect additional information about that particular toss, updating our credence about its result. SB is a form of such measurement, in which the intrinsic global probability and the credence assigned to a particular case can differ.
For example, if a Heads toss produced a louder sound, that could provide additional information about the result and raise our credence in Heads to, say, 0.6.
A useful term here is observation-selection effects, which may be clearer than “anthropics.” Some results are more likely to be observed. SB is a classical example: Tails is more likely to be observed.
The coin-in-a-crowd thought experiment
Suppose I toss a coin. If it lands Heads, I tell one person in a crowded room; if it lands Tails, I tell two people. A person who is told about the toss should then assign a 2⁄3 probability to Tails for that particular toss.
This is not an exact analogue of SB because the other members of the crowd continue to exist. Those who were not told can still assign probability 1⁄2 to Tails.
3. The no-MWI assumption
SB assumes that MWI is false, or at least that there are no other indistinguishable copies of Beauty. Thus, the alternative Heads branch does not exist when the result is Tails.
This assumption is necessary if SB is to support SIA through the claim that the Tails universe contains more observers.
However, if SIA implies that we live in the largest possible universe, and therefore in an MWI universe, then SB cannot serve as a test of SIA. This negative circularity weakens attempts to prove SIA through SB.
4. The sequential nature of events under Tails
Under Tails, Monday–Tails and Tuesday–Tails are not independent events, as explained in [ape-in-the-coat’s solution].
Beauty can predict her own actions on the other betting day because an exact counterpart of her exists there. This is the basis of the unusual two-thirder approach, discussed below, in which Beauty counts not only her own bets but also those of her exact counterparts.
5. There is no factual or testable uncertainty in SB
Once we specify a payoff matrix and agree on a global betting scheme, including whether the awakenings are altruistic toward one another, we can calculate a winning strategy for Beauty. The remaining uncertainty is interpretational.
Double-Extreme Sleeping Beauty shows that the choice can still matter when the stakes are extreme, and I do not know which answer is best. In our world, however, we may be more likely to inhabit a very large universe in which functionally similar experiments recur. Arguments such as the Presumptuous Philosopher may therefore fail to settle the issue.
6. SB will not help us escape the Doomsday Argument
The main practical application of SB is the Doomsday Argument. If Beauty learns that today is Monday, what should she believe about the probability that the coin landed Heads?
A halfer first assigns 1⁄2 probability to Tails and then divides that probability equally between Monday and Tuesday. This gives Monday–Tails a probability of 1⁄4, while Monday–Heads remains at 1⁄2. After renormalization, Beauty assigns Monday–Heads a probability of 2⁄3, suggesting a short timeline and an early Doom.
Rejecting the halfer position in SB does not eliminate the Doomsday Argument. Gott’s version is based on the idea that I am located near the middle of the total sequence of observers and does not explicitly compare short and long possible worlds. There is also Katja Grace’s SIA Doomsday Argument.
7. Beauty’s expectations about repetition
Beauty’s position depends on her expectations about whether this experiment, or a functionally similar one, will recur with her or with a functionally similar mind.
If the SB experiment represents a typical situation, she can act as though it is one member of a series, which supports thirdism. If the experiment will never recur, the information it provides has little value because the situation is unique. If it is typical, Beauty will reason from that typicality.
8. The two-thirder model
By symmetry with the double-halfer view, we can propose a two-thirder model. It begins by assigning equal probability to the three awakening-locations, giving Tails a total probability of 2⁄3. Beauty then does not update upon learning that today is Monday, because Monday–Tails being true always implies that Tuesday–Tails is also true. The long world therefore remains more probable.
This yields an anti-doomsday argument: I am more likely to inhabit a civilization that reaches an indefinitely long future.
The two-thirder model is an analogue of the double-halfer model for thirdism. In the double-halfer view, Beauty is a halfer but does not update upon learning that it is Monday, so her credence in Tails remains 1⁄2. In the two-thirder view, she likewise does not update, and her credence remains 2⁄3.
Bostrom’s hybrid model is a double-halfer model in the one-shot case but becomes a thirder model under repetition because it counts only actual agent-parts rather than merely possible ones.
Conclusion
One can argue that if we cannot solve SB, our minds are not capable of solving still more complex tasks such as AI safety. Yet we are also consuming a great deal of excellent researchers’ time on SB.
Perhaps SB is merely an intellectual game whose winner is the cleverest person who has ever lived. But there are many such games. Perhaps we hope to escape Doomsday predictions by defending thirdism and SIA. Yet SIA has a Doomsday Argument of its own.
SB may nevertheless help us analyze indexical problems for AI, such as how an AI should count other instances or copies of itself in simulations or parallel worlds.
Writers tend to become attached to their preferred solution to SB, leaving little room for uncertainty. I think we should accept our theoretical uncertainty, as I argued in the [Meta-Doomsday Argument](https://philarchive.org/rec/TURAMA-4).
There is also a view that any thought experiment can be replaced by a proper proof. What exactly do we want to prove with Sleeping Beauty? That possible copies should count as real?
I do not think SB is solved, because I am still uncertain what to do in Double-Extreme Sleeping Beauty. It will remain unsolved as long as reasonable people continue to disagree. But do we have any comparable logical or philosophical questions that are genuinely solved?
SB has one more hidden complication: a symmetry in which both the coin probability and the probability of being on Monday or Tuesday are 1⁄2. This symmetry can cause updates from SIA to cancel one another. Changing the coin’s probability or the number of copies breaks the symmetry.
I generated a detailed explanation of this cancellation with AI. Rather than include it here, I will put a link in the comments.
I think the Sleeping Beauty problem demonstrates that the concept of probability is not as solid as it seems, at least when we test it against unintuitive scenarios. If we ask how SB should bet if she’s trying to maximise expected payoffs of the whole experiment, there will be no disagreement. It’s only when we ask about “probability of heads” then the troubles start. Taboo “probability” and everyone will agree, wouldn’t they?
Yes, I think it is true. But there was a great post from Astral Codex Ten, which shows that it still maybe needed. There is also a post from ape_in_the_coat which shows that betting alone is also bad. I hope to address it in the next post which is now in draft.
There’s also the Solomonoff Induction location issue, where larger worlds get penalized if it takes many bits to locate observers within them.
Do you think it penalize tails?
In the standard SB question, not significantly, because the scenario is so contrived and the number of people under tails isn’t that high.
But in versions with thousands of people, yes.
Do we have a numerical estimate how much is the shift? ( of cause I can ask AI but interesting in your opinion). Something like 1/log N ?
Well a world with N people, with N large and not simple like a power of 2 or something, is going to be log2 N extra bits to specify over the same world with one person. Then an extra log2 N bits to locate a specific observer.
But since you might be any one of the N people, the second one cancels out and you’re left with the inherent complexity of the world. So log2 N, but if N is a simple number it can be much less. For N power of 2, it can be log log N. And you can have stuff that grows much faster than powers of 2, eg. N=Graham’s number doesn’t take many bits to specify.
The Sleeping Beauty Problem is mis-characterized. It is not a single probability experiment, it is one sample from a series of experiments. Without the amnesia drug, it would be just one experiment since the sampling would be redundant. But with the drug, each sample becomes independent sample. The unique circumstances of the problem conspire to hide that.
The way to resolve this, is to expand the problem. And before you claim that I am changing something, I am not. I am extending the exact same mechanisms in order to make their effects more intuitive.
Keep the amnesia. Instead of a coin, use a six-sided die. Instead of two days, make it six as well. And add a second “waking” scenario. Wake Beauty in either the morning, or the afternoon, and put a clock on her wall so she knows which kind of waking it is. The experiment will be governed by a 6x6 calendar, where the rows indicate the die roll, the columns indicate the day, and each entry in the calendar is either an “A” (for AM waking), a “P” (for PM waking), or an “S” (for sleep). And each time Beauty is awake, she is asked to assign a probability to each of the die results in {1,2,3,4,5,6} based on the time.
For my first version, each row must contain each letter at least once, but the remaining 18 entries are random. Here’s an example:
A P S S P P
A A P S S P
S A A P S P
A A S A P S
A S A S A P
P S P A P A
The Halfer argument is that Beauty gains no new evidence when she is awake, since she knew she would be awake during a morning, and an afternoon, at some time regardless of the die. Regardless of the clock, all of the probabilities should be 1⁄6. The Thirder Argument is that, if the clock says it is the afternoon, the probability that die rolled the number D is the number of times the letter “P” appears in row D, divided by the number of times it appears in the table; 12, in this example. So the probabilities are {3/12, 2⁄12, 2⁄12, 1⁄12, 1⁄12, 3⁄12}.
But what if we remove the restriction that each letter must appear in each row?
A S A P P S
S A A S P P
A S A A S A
S A A A A A
A P P S P S
P S S P S P
The same Thirder argument can be applied to this. If the clock says it is morning, a Thirder Beauty would say the answer is {2/14, 2⁄14, 4⁄14, 5⁄14, 1⁄14, 0}. But the Halfer argument falls apart. Either the answer is that there is still a 1⁄6 chance that the die rolled a six, even though beauty cannot be awake in the morning if the die rolled a six. Or the answer is {1/5, 1⁄5, 1⁄5, 1⁄5, 1⁄5, 0) and Beauty has to admit that what she observes does constitute new evidence.
Finally, what if, when Beauty is awake, they tell her that the afternoon technician became sick and every “P” was replaced with an “S”, like this:
A S A S S S
S A A S S S
A S A A S A
S A A A A A
A S S S S S
S S S S S S
This can’t change her second answer, regardless of whether she is a Halfer or Thirder. In other words, what Beauty bases her answers on, when she sees it is an AM waking, is whether each enrtry matches the evidence she has, or does not. Not what she would observe when it does not match.
Here is the crux of the Halfer’s anthropic solution. That an S in a cell not only removes that cell from Beauty’s sample space, it actually removes it from the experiment (i.e., from reality). And so the different numbers of “A”s in a row have to be collapsed into outcome, to make each row with an A have the same number—one—of “A”s.
Great example! I tried to get the same result with biased coin.
I think what Halfer misses is that probability of coin P is not the same that probability that the last toss of the coin was P. We can collect partial information about the last toss via different ways including observation selection effects. We can replace coin in your example with Sun exploding in the past and destroying earth with some chance. In that case, a priopi P of explosion is not the same that there was explosion but we survived (the example needs some polishing).
I’m not sure people disagree:
Given [her waking up], not knowing what day it is, she’s in the tails world 2⁄3 of the time.
Given [we’re about to play this game], there’s a 50⁄50 chance of heads or tails. So whatever day she wakes up on gives you no further information.
Does anyone disagree?
So 162 now? :-)
I submit that the whole complex of ideas around anthropics and decision theory is a mind killer.
I don’t think so. If we use simple anthropics, we can get useful predictions like duration of a process which we observe in a random moment in time. The whole trick to get sane anthropics is to exclude possible worlds and treat all worlds as real but with different densities.