Note that the math isn't showing up for me, might be a formatting issue.
There's a fatal flaw in the logic here.
Note that the martingale is defined over observations and expected observations.
But then you talk about the prior probability of heads. This is not meaningful, since she does not observe the coin flip.
Here's a consistent probability assignment that meets the martingale criteria.
On Sunday, she estimates probability ~1 of waking up tomorrow in a room. She doesn't have a probability on a particular room or of a coin result as of Monday because she doesn't observe that on Monday. She further estimates a probability of 2/3 for Tuesday room 1 and 1/3 for Tuesday room 2.
On Monday, she confirms that she woke up. She continues to have the same 2/3 and 1/3 probability assignment across the two rooms possibilities for Tuesday, since nothing has changed.
Alternatively, she can consistently maintain that each room is 1/2 likely.
I've got a different scenario with no duplicates. A fair coin is flipped today, if it's heads, Roulette Rory gets killed with 75% probability before learning what the coin says. If it's tails Rory survives. We ask Rory what her probability is. She says today her probability is 50/50, but if she learns about results she is 80% confident in tails.
Why is it ok for Rory to expect her future self to be 80% confident in something that she's currently 50% confident in, but it's not ok for Beauty to do the same? In both cases, the number of observers is varied in a way that's correlated with the observation. That's all that's happening here.
Sorry, the latex display should be fixed now.
She doesn't have a probability on a particular room or of a coin result as of Monday because she doesn't observe that on Monday.
Pertinent comment, so I've added an
I've got a different scenario with no duplicates. A fair coin is flipped today, if it's heads, Roulette Rory gets killed with 75% probability before learning what the coin says.
In this case, not making an observation (due to death) is another possible history. The possible future histories are "see heads", "see tails", and "{}" (see the other post on how deaths don't affect Bayesian probability).
That's also the reason why ID(Sleeping Beauty) doesn't have the paradox; the duplicates existing but not observing their history means that heads Sleeping Beauty gains information on Monday (I'm not dead, hence I wasn't the duplicate that died).
I don't think the epsilon chance of observing now really helps. It makes the current probabilities meaningful but you still have no particular reason to insist that the current probabilities ("what happens if I were to observe this today") matches the future probabilities ("what happens if I observe it in the future"), given that the observations are correlated.
Re death, it just doesn't seem different to me. If I can magically increase my probability in outcome A by committing to suicide in not-A worlds, I can also increase my probability in A by duplicating myself in A worlds. You've set up a formalism that allows the former and not the latter, then you're pointing to a contradiction. But they're both about as mysterious and settled the same way. The first isn't controversial and the second shouldn't be either.
(The exact probabilities still depend on SSA vs SIA vs others. But the fact that they won't match the objective probabilities seems fine.)
If I can magically increase my probability in outcome A by committing to suicide in not-A worlds
You can't. You can increase your relative probability of seeing outcome A, conditional on you being alive. You split into: see A, see not A, be dead. Both of the later are in world not A, and you can shift their relative size, but not their total size.
It's survivorship bias via suicide, nothing strange about it.
Any conditional duplication is (for the purposes of the scenario's credence of heads when asked) identical to unconditionally duplicating and then killing the duplicate in room 2 on heads before awakening.
What's more, it's the same as just not asking the duplicate in room 2 on heads, without any killing involved.
What's more, it doesn't matter when that duplication happens prior to being asked, before the coin flip or after, so longer as neither knows which one is going to be asked on heads.
So the duplication isn't the problem.
SIA (which I tend to agree with) treats unconditional and conditional duplication as the same; but that's an assumption we don't get to take for free.
On heads, being killed or not being asked have clear effects in terms of expectation of alternative observations to see or not see.
That's why unconditional duplication is fully martingale, while conditional duplication isn't.
I'm not sure whether you're agreeing with my first point or not, regarding the equivalence of conditional duplication with unconditional duplication + conditional killing.
Since you said that the martingale condition is fully satisfied for both unconditional duplication and for conditional killing, do you believe that it is necessarily (regardless of SIA belief) satisfied for the sequence of both?
If not, how does the first equivalence break?
Yep, the martingale is satisfied for both.
It's the equivalence between (conditional duplication) and (unconditional duplication + conditional killing) that I'm disputing (or at least saying that you can't just assume for free).
The difference is clear for the formal martingale: let's count the full agent histories from beginning to end. Conditional duplication has three histories - (Sunday, awake, room 1, T), (Sunday, awake, room 2, T) and (Sunday, awake, room 1, H). While (unconditional duplication + conditional killing) has four - those three plus extra one heads: the "death" history of just (Sunday).
Thus the martingale in the (ud + ck) case has an extra possible future observation to consider: specifically, the non-observation. That's why "awake" gives information in the (ud+ck) case (not all your duplicates would be guaranteed to see it) but not in the (cd) case (where all your duplicates are guaranteed to see it).
(It doesn't help if one says that "non-observations don't count", because then the theory breaks breaks the martingale in the case of just "conditional killing" on its own.)
Now, philosophically, I'm partial to the idea that "instant duplication followed by deletion" should be the same as "not creating duplicates at all"; that's a strong pro-SIA argument. But that doesn't fix the martingale; it instead shows that if we assumed that there were instantly-deleted duplicates, then that assumption would fix the martingale.
I consider the cd = ud+ck equivalence to be much more fundamental than any martingale property, so I guess that's the difference.
If we have an epistemic model in which we need to consider whether or not a duplicate was created and destroyed without ever making any observations, then as I see it that model should be discarded as it contains internal flaws. It shouldn't even make any difference if we bring in an inert lump of random matter instead of a non-conscious duplicate, either. From an epistemic point of view, all non-observers are equivalently irrelevant.
Sure, but then why is it strange to do the same with duplication? You're shifting your relative probability of seeing A, by creating extra versions of you that see A.
Predicted observation. You know you'll be observing waking up on Monday (or, at least, that every thread going on from you to will observe waking up on Monday). And yet you are not, now, currently in the epistemic state you will be on Monday, despite know exactly what you will observe.
Bbtw, death and birth/creating are opposites, and neither is really a problem in anthropic probability. The opposite of duplication is merger - much more difficult to achieve, but, yes, it has similar issues to duplication.
"The thread of my experience has forked/merged" is something probability theories were not built to deal with; unlike threads of experience ending and starting, which is can deal with fine.
I don't buy this distinction.
Every thread going on from you will observe something in Roulette too. (The threads that die don't go on from you, then.)
When predicting your future experiences, you automatically condition on the fact that there will be future experiences, so they won't always match objective probabilities. And you can also adjust for some versions of you having more experiences than others.
It's more fragile than that. E.g. suppose a coin is flipped and on tails, you'll be given a false memory of being asked your credence for heads, and answering "100%". On heads you'll simply be asked the question with no memory alterations. Then in both cases you'll be asked "are you sure?"
If you're in the situation of remembering having been asked the question but not answering, then you must be in the heads case, and should answer "100%". But then when you're asked "are you sure", you are no longer sure because you now have exactly the same sort of memories that would have had in the tails case. Your memory hasn't been altered in this case and you have no new information, so why are you now uncertain about something you were previously rationally certain of?
Now let’s consider Monday. Going to sleep on Sunday, Sleeping Beauty knew she would awaken, so martingale guarantees that the probability can’t change on Monday
Here's why I think violating Martingale here is fine: There's no Dutch book against this violation. If Beauty bets on Sunday in favor of Heads and on Monday in favor of Tails, the Monday bet happens twice if Tails. (This is analogous to ordinary Sleeping Beauty, where an "awakening bet" happens on both Monday and Tuesday given Tails.)
If people think Dutch book resistance is a good argument for Bayesianism (and Martingale updating in the usual case), they should find the Dutch book resistance of SIA / generalized thirding to be a compelling reason to deviate from Martingale.
I imagine a conversation between a Thirder and a Bayesian as follows:
Bayesian: "Thirding violates Martingale. That's terrible, and not Bayesian at all."
Thirder: "If that's true, what's so great about being Bayesian?"
Bayesian: "Dutch book arguments show that in non-amnesic scenarios, Martingale and Bayesian updates are uniquely resistant to Dutch books."
Thirder: "But by various theorems, Thirding is uniquely resistant to Dutch books in amnesic situations."
Bayesian: "..."
(and similar could go for duplicate scenarios, not just amnesic scenarios)
See my recent post for more on SSSA+SIA resistance to Dutch books.
Yep, I fully agree :-) I've already said that decision theory should prime over probability theory; examples like this are why.
Any probability theory is incomplete, though, because if there's a Dutch book where the people paying and receiving are arguably different people, the Dutch books outcome shifts depending on how they value each other. So Sleeping Beauty who is an average utilitarian over her future copies will behave like SSA has the true probabilities (for selfish bets).
If you ask about probabilities of observer-events not relying on personal identifications, none of this is any problem whatsoever. There are epistemic symmetries between some observer-events, but not all.
In particular, the martingale condition relies on the fact that ordinarily there is a certain type of time-translation symmetry over personal identity. In this case there obviously isn't, so that applying the consequences of such a symmetry is simply incorrect.
Ok! So you agree with the impossibility result, just disagree that martingale is a reasonable property in the first place.
That's fine; just be aware of what you might be giving up. Which is that now you must update on entirely predicted observations. e.g. if duplication can happen in the middle of a conscious experience (so no sleeping required) then you might know that, for instance, you will update in the middle of an exercise routine.
Yes, I have posted an example of suddenly changing rational updates with nothing unusual happening in these comments, and previously elsewhere.
As an aside, note that decision-theory-with-precommitments has no problem managing duplication events. That’s why I consider decision theory as the more fundamental object; probability theory is a subset of decision theory where the utility function is the Brier score or any other strictly proper scoring rule.
The utility function will also encode how to aggregate Brier scores across multiple duplicates; summing will lead to SIA-like behaviour, while averaging leads to SSA-like behaviour. In that view, a Sleeping Beauty who sums Brier scores and chooses to create a billion copies is not changing her current or future credence of tails; instead, she’s maximising her utility in the tails world.
This is also my takeaway from stuff like OSAC and Non-Realist UDASSA, where probabilities in general are much more like caring measures/utility functions than they are beliefs, and a caring measure expresses how much the agent cares about how things go in particular worlds.
This is also why contra David Matolcsi, I don't find the subjectiveness of his solutions to be an actual problem.
Put another way, utility theory/decision theory is more general than probability theory, and the goal in those anthropic problems/acausal trade problems/decision theory problems is not to get a correct model according to a specific model of the world, but rather to maximize utility in various worlds.
This is unavoidably, irretrievably subjective, but you can learn to at least accept it as a basic precondition for doing well at these problems, even if you aesthetically dislike it.
Isn't your application of Bayes's rule wrong?
P(T | awake, room 1) = P(room 1 | T, awake) P (T | awake) / P(room 1 | awake)
The second and third factors on the RHS are exactly the values under contention.
No, that's just Bayes, conditional on another factor, awake (I should have written it last for higher clarity).
Bayes is: P(A|B)=P(B|A)*P(A)/P(B)
Bayes with an extra condition is: P(A|B,C)=P(B|A,C)*P(A|C)/P(B|C)
(or you can define Q(X)=P(X|C) and do standard Bayes with Q(X))
The problem is that we don't know what P(T | awake) is (the P(A|C) factor in your notation). That's the whole question at hand. You are implicitly assuming some value for it in this step.
I am applying a theory of probability that I'm taking as reasonable until I reach a contradiction that shows that it isn't (proof by contradiction). If P is reasonable, it should give a value to statements like P(T | awake), and that value should follow the normal Bayes rules.
I don't think I ever injected a specific value that wasn't already there in the assumptions?
You say: "being in room 1 is equally consistent with
Here it is with plugging in an unknown t instead of assuming a value of Q(T|awake):
Q(T| awake, room 1) = Q(room 1 | T, awake) Q (T | awake) / Q(room 1 | awake)
If Q(T|awake) = t, then Q(H | awake) = 1-t and Q(room 1| T, awake) = 0.5 etc. give that Q(room 1 | awake) = t /2 + (1-t) = 1 - t/2
So:
Q(T| awake, room 1) = (1/2) (t) / (1-t/2) = t/(2 - t)
For this to equal your value of 0.5, you must be assuming that t := Q(T|awake) = 2/3. I suspect you heuristically did "room 1 is 'equally consistent' with T and H meaning the the likelihoods are equal and so their posteriors are equal". However, the likelihoods are not the same: room 1 is more consistent with H than with T in the sense of likelihoods and Bayes rule uses this. If you still think I'm wrong, please write out exactly the computation you're claiming, which you haven't yet done.
To get the values of Q(T| awake, room 2) and Q(T| awake, room 1), I am applying "simple Bayes" because upon seeing awake and room 1, there is no longer any anthropic uncertainty as to who you are in the world.
Then, formally, I am assuming that Q is a well defined probability in general (obeying all the Bayesian equations, amongst others) and aiming to get a contradiction from this assumption.
So I set up an equation involving terms like Q(room 1 | awake) in order to infer what these would have to be. I've used simple Bayes to pin down some of the terms, and then I'm using the assumption that Q is well defined to set up an equation between the known terms and the unknown terms.
So one can write:
Q(T| awake, room 1) = (1/2) (t) / (1-t/2) = t/(2 - t)
And then, plugging in the known value of Q(T | awake, room 1), one deduces Q(T | awake) = 2/3. I haven't assumed that value; I've deduced it, under the "Q is reasonable, martingale, and simple Bayes works" assumptions (actually, all that I can deduce - and all that I need - is that Q(T | awake) > 1/2; getting it to be 2/3 requires slightly stronger symmetry assumptions).
I can also deduce Q(T | awake) = 1/2 (martingale from Sunday to Monday), and thus get a contradiction. One of my assumptions must be wrong. Q cannot simultaneously obey Bayesian updating, obey the martingale, and obey simple Bayes.
"plugging in the known value of Q(T | awake, room 1),"
But that is not known! That's the value I am contesting your derivation of. You have asserted that it is 0.5 supposedly due to Bayes law but I am telling you that the math does not work out that way. If you want to convince me that Bayes law forces Q(T|awake, room 1) to be 0.5, then you need to write out that application of Bayes law and derive the number 0.5, not simply plug it in.
Ok, got the issue. Thanks!
The condition I'm calling "simple Bayes" is that, if there is no doubt as to which agent you are in the world, then proceed by taking the prior over worlds and updating on the evidence "there exists a person in this world who have made this observation" (in non-anthropic situations, "there exists a person who has observed history H" and "I have observed history H" contain the same information).
On Monday, there is uncertainty are to which agent SB is. On Tuesday there is not: the two agents in the Tails world can tell each other apart, based on room number observation.
So simple Bayes applies to Sunday and Tuesday, but not to Monday.
On Sunday, the observations are independent of the coin flip, so P("I saw Sunday" | w_T) = P("I saw Sunday" | w_H) = 1.
If h_0 = ("I saw Sunday"), then P(w_T | h_0) = P(h_0 | w_T)P(w_T)/((P(h_0 | w_T)P(w_T) + P(h_0 | w_H)P(w_H)) = 1*(1/2)/(1/2+1/2)=1/2. Same thing for P(w_H | h_0).
On Tuesday, in the tails world, there exists, with certainty, an agent who has observed: h_1=("it's Sunday", "I'm awake on Monday", "it's Tuesday and my room number is 1"). Similarly, in the heads world, there exists, with certainty, an agent (the only agent) with the same observation.
Since both exist with certainty, then the formula is the same as before: P(w_T | h_1) = P(h_1 | w_T)P(w_T)/((P(h_1 | w_T)P(w_T) + P(h_1 | w_H)P(w_H)) = 1*(1/2)/(1/2+1/2)=1/2, and the same for P(w_H | h_1).
So the naive direct application of Bayes in non-anthropic situations gives us these values. The impossibility result is just that, given these values for Sunday and Tuesday, there are no values on Monday (the time-slice where we can't use simple Bayes because there is genuine uncertainty as to which agent SB is) that allow the martingale condition to extend between Sunday and Tuesday.
Note that I'm not saying that SIA or SSA are wrong or that you can't do anthropic probability. I'm saying that if you do do anthropic probability, you have to drop some intuitive properties possessed by standard probability.
For instance SIA drops the martingale condition from Sunday to Monday (indeed SIA always obeys simple Bayes). SSA (which is less uniquely defined) either drops the martingale condition from Monday to Tuesday or drops simple Bayes on Tuesday (using "centered worlds" is an explicit acknowledgement of dropping simple Bayes).
Thank you for the explanation. We need some better notation for this: write I(x) to mean that I observed x (or will observe it). Write E(x) to mean that I know that there exists with certainty someone who observed x (or will observe it).
First, in general, we don't expect P(I(x)) = P(E(x)). For example, P(I(room 1)|I(room 2)) = 0 but P(E(room 1) | I(room 2)) = 1.
You define "simple Bayes" to mean:
P(y | I(x)) = P(y | E(x)) for any y and x
I would argue that you really need to pick a totally different name for this property since it doesn't have anything to do with Bayes' law. (If anything, I would call it "nonstandard Bayes".) And your definition of it in your article needs to be clearer: as stated, it's just about I(x) without mentioning E(x).
This is also the same crux that the paradox has always revolved around: does finding out who you are give you information? I think that's closer to the standard phrasing and makes it more clear what you're being asked to give up or not. You're not being asked to give up Bayes' law: that's always true.
So we can apply simple Bayes again. Being in room 2 compels
, while being in room 1 is equally consistent with or . So while .
I don't understand your justification for why
Could you elaborate?
According to my calculations
Being in room 2 compels
Did you mean
I have had an intuition of my own for why
Before, it seemed correct that
But now I believe that
Why do I believe this?
The martingale condition/conservation of expected evidence applies only if the relevant epistemic state is exactly the same before and after observing the evidence. By "epistemic state" I mean the set of statements the agent knows is true. On Sunday SB knows she is the original, on Monday she does not. So on Sunday her epistemic state contains the statement "I am the original", on Monday it does not. This change of epistemic state before observing evidence "awake" and after observing evidence "awake" doesn't allow the application of the martingale as you described. Because if she learns she is NOT the original, then her probabilities for Heads and Tails differ from the priors. For the martingale to apply, additionally she has to know on Sunday that she will know she is the original on Monday. Which is not the case.
According to my calculations
From your second post:
Based on the above, I believe the correct math would be:
So on Sunday her epistemic state contains the statement "I am the original", on Monday it does not.
Epistemic states can change freely without breaking martingales; on Sunday, she's in room 1; put her to sleep and wake her up in room 1 or 2, randomly chosen; no duplications. Her epistemic state has moved from "I'm in room 1" to "I don't know what room I'm in", and yet the martingale works perfectly across this.
In any case, the statement is not "I am the original" but "on Monday, I was the original", which is preserved.
I've also weakened the requirements; now it's the law of total probability that is violated, not the martingale.
If you get
I realized some of my mistakes since my previous post.
Not to get into overmuch details, I'll just say I'm back at where I started.
I believe:
In the context of your second post:
If "simple Bayes" implies that upon learning she is the original on Tuesday SB must have the same posteriors as she had priors on Sunday when she also knew she was original, then I don't see why should I accept "simple Bayes" as a property of probability.
If she learns on Tuesday "I am the original" is false, then her probability for Tails becomes 1, which makes this positive evidence for Tails. Therefore, learning on Tuesday "I am the original" is true must be negative evidence for Tails. If she doesn't lower her probability for Tails upon learning she is the original, that would be violation of conservation of expected evidence. In other words, confirmation bias - no matter what she learns on Tuesday (out of available options), her probability for Tails may go up but never down.
Simple Bayes means that, if there is no longer any anthropic angle to your problem, you should be able to apply non-anthropic updating.
Equivalently-ish [1] it implies that it doesn't matter when you started thinking anthropically; if you started doing it yesterday and then updated on one day's observation, it would be the same as if you started doing it today, or twenty years ago and updated on twenty years of observations. And it also doesn't matter that you could have been X, if you know that you're not X.
Formally, this is anchor-free: there isn't an anchor moment that defines the beginning of when you're thinking anthropically [2] .
Anchor-free is stronger than simple Bayes; simple Bayes is just the equivalent of anchor-free in situations where there is no anthropic uncertainty. Anchor-free is the real condition I want, but simple Bayes is enough for the impossibility result. ↩︎
One needs to be careful here; technically, any method with a specified anchor is "anchor-free" in a certain sense if it says "whenever you actually started thinking anthropically, then go back to this moment in the past (or future) and anchor then". If the specified anchor is part of the definition of the method, then it's "anchor-free" in that it formally doesn't matter when you started thinking anthropically, because in all cases you just go to a specified anchor [3] and role out everything backwards/forwards from that I don't consider "anchor specified" to be anchor-free. ↩︎
"Go to a specified anchor" is not actually that trivial to define. If the "anchor specified" truly behaves like "anchor-free" then it has to work from the very moment of first consciousness (as that's certainly a place you could start anchoring from), when you could have been, potentially, any being in the universe. So there's a universal anthropic probability that depends potentially on the details of all agents who ever were or ever could be. ↩︎
Yes, I did mean
Will link to a new post that explains the impossibility result better.
tl;dr
In this post, I’ll extend beyond basic anthropic problems and see what happens when duplicates or copies are allowed. I’ll start with an impossibility result that may help clear up some of the confusion in anthropic probability: namely that no reasonable probability theory can stay consistent across duplication events.
It’s the duplication event that is the issue. The presence of duplicates or copies is not a problem. Giving birth or creating new agents is not a problem. But duplicating an already existing agent breaks anthropic probability.
That being said, just because no anthropic probability theory is perfect, doesn’t mean that some aren’t better than others. SIA is a top candidate for an anthropic probability theory (and indeed it is consistent before and after duplication events).
In the final post in the series, we’ll see some of the issues with SIA and infinity, and I’ll introduce D-SIA, distributional SIA, which fixes many of the issues.
Inconsistency of probability across duplication
To illustrate, I’ll be using yet another variant of the Sleeping Beauty problem.
In the traditional problem, a coin is secretly tossed, and Sleeping Beauty is put to sleep on Sunday. She reawakens on Monday. If the coin was heads, that’s the end of the experiment. If the coin is tails, she is fed a sleeping draught with an amnesia potion that erases her memory of Monday, and is then reawakened on Tuesday, initially unable to distinguish the Monday awakening from the Tuesday awakening.
The incubator variant has no initial Sleeping Beauty and instead has two identical rooms, 1 and 2. A coin is tossed, and on heads, one Sleeping Beauty is created in room 1, and on tails, two identical Sleeping Beauties are created, one in each room.
The original Sleeping Beauty problem has memory loss, which you can argue is a loss of rationality or confusion about their past experiences. The incubator variant has no initial Sleeping Beauty. So I’ll be using a new variant, Duplicate Sleeping Beauty.
Duplicate Sleeping Beauty: there is an initial Sleeping Beauty and two identical rooms. Sleeping Beauty is put to sleep on Sunday and a coin is tossed. If it’s heads, she is put in room 1 and awakened on Monday. If it’s tails, she is duplicated, and one duplicate is placed in each room, and they are both awakened. On Tuesday, she will be informed of which room she is in.
Then I’d claim:
First, let’s define “reasonable”.
Martingale and simple Bayes
One of the key properties of a “reasonable” probability function is that it has the martingale condition (also known here as conservation of expected evidence). Specifically:
Suppose I asked for the probability of future observation happening (versus not happening). And suppose that I asked for the probability that my full future history contains the observation (versus all my other future histories).
The martingale property of probability says that these two questions are the same. It’s a fundamental property of probability, and it’s hard to see what it would mean for the two to be divergent. Note that the martingale is defined over observations and expected observations.
The other reasonable property is simple Bayes: if an agent has no uncertainty about who they are in the world, then the agent’s probability can be derived by Bayesian updating of their observations on the prior.
Constructing the impossibility
In Duplicate Sleeping Beauty, there are three moments to consider: Sunday, before Sleeping Beauty is put to sleep. Monday, after she is awoken. And Tuesday, after she is informed of which room she is in.
Then there are no identical duplicates on Sunday or Tuesday, so simple Bayes can be used to give probabilities. Then we’ll show there is no probability assignment on Monday that is martingale-consistent with Sunday and Tuesday.
Specifically, let and be the worlds where the coin is heads ( ) or tails ( ), respectively. The prior is . Let be the probability function we are attempting to fit to the whole setup.
Now on Sunday, there is a single agent, no observations so far, so Bayesian updating gives and (since we defined the martingale for observations, not events, we need to add an chance that Sleeping Beauty will get to observe the coin result on Sunday).
On Tuesday, the agent has awoken and been told which room she is in. So each agent has seen (awake, room 1) or (awake, room 2). There is one agent in , who has seen (awake, room 1). There are two agents in , but they have each seen different things: there are no longer any indistinguishable duplicates.
So we can apply simple Bayes again. Being in room 2 compels , while being in room 1 is equally consistent with or . So while .
Now let’s consider Monday. Going to sleep on Sunday, Sleeping Beauty knew she would awaken, so martingale guarantees that the probability can’t change on Monday: .
Then the martingale condition gives:
Thus .
I’ll now add another technical probability requirement: that, since (awake, room 2) is actually possible, then is greater than zero. This gives our contradiction.
Creation and birth versus duplication and merging
The problem only arises because of the duplication event. If a different Sleeping Beauty is created ex nihilo in room 2, that’s not a problem. If Sleeping Beauty gives birth overnight to a new agent in room 2, that’s also not a problem. The problem is that the subjective timeline of a single well-defined agent has split, and that messes everything up.
It’s less physically plausible, but if the two copies were to merge, that would mess things up in the other direction. Agents being created is not a problem; agents dying is not a problem. It’s the forking and merging of timelines that is the problem.
Anthropic disputations
Pause for a moment to sympathise with the various people who have analysed and disputed anthropic probability. No matter how obvious it seems to you that one or another option is right, it’s clear that something very weird is going on. You might think of many “halfers” as people who took the initial 50-50 probability of the coin toss and tried to push it forward with the martingale condition. And you might think of many “thirders” as people who took the final probabilities when Sleeping Beauty knows which room she is in, and tried to push that backward with the martingale condition.
Spare also a thought for those who tried to insist that, actually, somehow, both the Sunday and the Tuesday simple Bayes were right and that something weird but fundamentally acceptable was happening on Monday. Basically, something is breaking down in the standard laws of probability, and people are trying to patch it.
SIA specific issues
When discussing that argument, someone made the point that, on Sunday, Sleeping Beauty knew exactly who she was, while on Monday, she had lost that knowledge; hence that just being awake Monday gave her some new evidence. Since it’s new knowledge, she should update on it: even if “being awake” is entirely predictable as an observation, she still enters a new epistemic state, and her probability should be allowed to reflect that.
I understand the argument but disagree. Contrast with a simpler situation: I’m in room 1 or room 2, no duplication or anthropic effect. I’ve seen “Room 1” painted on the door, so I know I’m (almost certainly) in room 1. However, tomorrow, I will discover that actually the painter mispainted, and painted room 1 on both rooms; thus, tomorrow, I will lose exact knowledge of which room I’m in.
What the martingale and conservation of expected evidence are saying is that, if I know today that tomorrow I will lose exact room knowledge... then I have actually lost that exact room knowledge already. The martingale breakdown is not that Sleeping Beauty, on Monday, won’t know who she is. It’s that, despite knowing this fact, today she does know who she is.
We can be more precise. Sunday Sleeping Beauty knows that the Monday Sleeping Beauties will lose knowledge of who exactly they are. And the Monday Sleeping Beauties know that too. So none of the agents are in any doubt as to the epistemic states of the other agents. The breakdown is that even though all agents agree on all facts about the universe and about each other’s knowledge and have the same priors, their probabilities differ.
The power to change the past
We can illustrate the probability breakdown further. Let’s assume that the martingale condition holds from Monday onwards, since there are no further duplications. Then, under mild conditions, Sleeping Beauty will follow SIA from Monday onwards[1].
Thus with her current actions, she can predictably and directionally change her future credence of a past event. That is not how standard probability works.
Decision theory
As an aside, note that decision-theory-with-precommitments has no problem managing duplication events. That’s why I consider decision theory as the more fundamental object; probability theory is a subset of decision theory where the utility function is the Brier score or any other strictly proper scoring rule.
The utility function will also encode how to aggregate Brier scores across multiple duplicates; summing will lead to SIA-like behaviour, while averaging leads to SSA-like behaviour. In that view, a Sleeping Beauty who sums Brier scores and chooses to create a billion copies is not changing her current or future credence of tails; instead, she’s maximising her utility in the tails world.
Defending SIA
Though SIA has problems around duplication events, it has no such problems after the event. It can deal with duplicates, copies, Sailor’s children, anthropic events involving risks of death or extinction, and similar.
So it is still, in my view, the best candidate for (non-infinite) anthropic probability. In as much as anthropic probability can be made to make sense, SIA seems to be the best candidate around.
How so? Well, I consider the martingale breaking down between Monday and Tuesday (where there are no duplication events) to be much worse than breaking down between Sunday and Monday (where there is a duplication event).
The most convincing argument to me is that in non-anthropic situations, SIA seems inarguably correct.
On top of that, there are modifications of duplication problems[3]which SIA handles perfectly fine.
In the next post, the last in the series, we’ll show the problems SIA has with infinity and how a variant, D-SIA (distributional SIA), can fix these.
Suppose that on the heads branch, instead of being put in room 1 with certainty, she is put in room 1 with probability . The first mild condition is that a) This can be modelled as a mix of two duplicate Sleeping Beauty problems: with probability , one where she is put in room 1 with certainty, and with probability , one where she is put in room 2 with certainty. The second mild condition is b) that those two problems are exactly the same, by symmetry, up to exchanging the room labels. ↩︎
See also “Defeating Dr. Evil with Self-Locating Belief”. ↩︎
Let be an anthropic probability problem involving duplication. We define , the inert duplication variant of by the following:
Then any reasonable theories of anthropic probability will agree with each other in, and the probabilities that they will give are the same as SIA in (some versions of SSA may require that the duplicates be allowed to run briefly before being deleted). Moreover, SIA does obey the martingale condition on .
What’s changed here? Think of Sleeping Beauty again. Because of the inert duplicate in the heads world, when she wakes up on Monday, she will downgrade the probability of the heads world. Why? Because she could have been the inert duplicate. So “waking up” now carries information: she is not an inert duplicate. The Sunday Sleeping Beauty expects that she will experience waking up on Monday or will experience nothing, so waking up does carry information. The “experience nothing” carries away exactly enough probability mass that everything is consistent from then on. And the Chosen world-size Sleeping Beauty? Here it’s not Sleeping Beauty changing the probability of past events. Instead it’s her playing a duplicate version of Quantum suicide: she’s sacrificing her duplicated existence in the heads world so that her surviving copy will find heads very unlikely.
So, though SIA can’t claim to be martingale across duplication events, it can claim to be martingale on similar setups that are at least arguably isomorphic.
Warning: people who are SIA fans may find this argument more convincing than it seems. An SSA fan could point out that SIA is unchanged by inert duplicates, while SSA is. Therefore, it’s not surprising that if we insert inert duplicates in the right places, we can get SSA to vary until it matches up with SIA. It’s interesting that all anthropic probability theories seem to match up at this point, but, they could say, SIA and SSA are the only serious contenders, so if they match up, everything does. ↩︎