# User: Stuart_Armstrong
Profile URL (HTML): [/users/stuart_armstrong](/users/stuart_armstrong)
Profile URL (Markdown): [/api/user/stuart_armstrong](/api/user/stuart_armstrong)
* Karma: 18431
* Alignment Forum karma: 1907
* Posts: 694
* Comments: 3757
* Member since: 2009-03-26 10:25:39Z
Bio
---
*No bio.*
Top Posts
---------
### [The AI in a box boxes you](/api/post/the-ai-in-a-box-boxes-you)
By [Stuart_Armstrong](/users/stuart_armstrong)
2010-02-02 10:10:12Z
* Karma: 179
* Tags: [AI Boxing (Containment)](/w/ai-boxing-containment), [Simulation Hypothesis](/w/simulation-hypothesis), [Anthropics](/w/anthropics), [Mindcrime](/w/mindcrime) (Frontpage)
Read more: [/api/post/the-ai-in-a-box-boxes-you](/api/post/the-ai-in-a-box-boxes-you)
### [Using GPT-Eliezer against ChatGPT Jailbreaking](/api/post/using-gpt-eliezer-against-chatgpt-jailbreaking)
By [Stuart_Armstrong](/users/stuart_armstrong) with [rgorman](/users/rgorman)
2022-12-06 19:54:54Z
* Karma: 170
* Tags: [Computer Security & Cryptography](/w/computer-security-and-cryptography), [GPT](/w/gpt), [AI](/w/ai) (Frontpage)
Read more: [/api/post/using-gpt-eliezer-against-chatgpt-jailbreaking](/api/post/using-gpt-eliezer-against-chatgpt-jailbreaking)
### [Assessing Kurzweil predictions about 2019: the results](/api/post/assessing-kurzweil-predictions-about-2019-the-results)
By [Stuart_Armstrong](/users/stuart_armstrong)
2020-05-06 13:36:18Z
* Karma: 152
* Curated
* Tags: [Forecasting & Prediction](/w/forecasting-and-prediction), [Forecasts (Specific Predictions)](/w/forecasts-specific-predictions), [World Modeling](/w/world-modeling) (Frontpage)
Read more: [/api/post/assessing-kurzweil-predictions-about-2019-the-results](/api/post/assessing-kurzweil-predictions-about-2019-the-results)
Recent Posts
------------
### [III. Anthropic reasoning has issues with infinite worlds; D-SIA can fix this](/api/post/iii-anthropic-reasoning-has-issues-with-infinite-worlds-d)
By [Stuart_Armstrong](/users/stuart_armstrong)
2026-08-03 21:31:53Z
* Karma: 11
* Tags: None (Frontpage)
Read more: [/api/post/iii-anthropic-reasoning-has-issues-with-infinite-worlds-d](/api/post/iii-anthropic-reasoning-has-issues-with-infinite-worlds-d)
### [II. Anthropic reasoning with duplication is not consistent with the usual probability properties](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent)
By [Stuart_Armstrong](/users/stuart_armstrong)
2026-08-03 14:21:23Z
* Karma: 5
* Tags: [World Modeling](/w/world-modeling) (Frontpage)
Read more: [/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent)
### [I. Anthropic reasoning without duplicates is just standard Bayesian updating](/api/post/i-anthropic-reasoning-without-duplicates-is-just-standard)
By [Stuart_Armstrong](/users/stuart_armstrong)
2026-07-30 13:08:29Z
* Karma: 29
* Tags: None (Frontpage)
Read more: [/api/post/i-anthropic-reasoning-without-duplicates-is-just-standard](/api/post/i-anthropic-reasoning-without-duplicates-is-just-standard)
### [Value Generalisation 3: Pre-aligned AIs](/api/post/value-generalisation-3-pre-aligned-ais)
By [Stuart_Armstrong](/users/stuart_armstrong)
2026-07-29 15:58:16Z
* Karma: 16
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/value-generalisation-3-pre-aligned-ais](/api/post/value-generalisation-3-pre-aligned-ais)
### [Value Generalisation 2: The Missing Hole in AIs’ abilities](/api/post/value-generalisation-2-the-missing-hole-in-ais-abilities)
By [Stuart_Armstrong](/users/stuart_armstrong)
2026-07-29 15:58:00Z
* Karma: 16
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/value-generalisation-2-the-missing-hole-in-ais-abilities](/api/post/value-generalisation-2-the-missing-hole-in-ais-abilities)
### [Value Generalisation 1: a Research and Deployment Program](/api/post/value-generalisation-1-a-research-and-deployment-program)
By [Stuart_Armstrong](/users/stuart_armstrong)
2026-07-29 15:57:50Z
* Karma: 22
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/value-generalisation-1-a-research-and-deployment-program](/api/post/value-generalisation-1-a-research-and-deployment-program)
### [The true "test" dataset for a generalised task](/api/post/the-true-test-dataset-for-a-generalised-task)
By [Stuart_Armstrong](/users/stuart_armstrong)
2026-07-27 16:16:03Z
* Karma: 23
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/the-true-test-dataset-for-a-generalised-task](/api/post/the-true-test-dataset-for-a-generalised-task)
### [Occam’s razor is about using the past to predict the future](/api/post/occam-s-razor-is-about-using-the-past-to-predict-the-future)
By [Stuart_Armstrong](/users/stuart_armstrong)
2026-07-15 19:35:11Z
* Karma: 52
* Tags: [Rationality](/w/rationality), [World Modeling](/w/world-modeling) (Frontpage)
Read more: [/api/post/occam-s-razor-is-about-using-the-past-to-predict-the-future](/api/post/occam-s-razor-is-about-using-the-past-to-predict-the-future)
### [Value generalisation: value correction](/api/post/value-generalisation-value-correction)
By [Stuart_Armstrong](/users/stuart_armstrong)
2026-07-10 07:56:04Z
* Karma: 25
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/value-generalisation-value-correction](/api/post/value-generalisation-value-correction)
### [Pragmatic FDT, and predictors as game theory](/api/post/pragmatic-fdt-and-predictors-as-game-theory-1)
By [Stuart_Armstrong](/users/stuart_armstrong)
2026-07-03 13:22:38Z
* Karma: 36
* Tags: [AI](/w/ai) (Frontpage)
Read more: [/api/post/pragmatic-fdt-and-predictors-as-game-theory-1](/api/post/pragmatic-fdt-and-predictors-as-game-theory-1)
Recent Comments
---------------
### Comment by [Stuart_Armstrong](/users/stuart_armstrong) on [II. Anthropic reasoning with duplication is not consistent with the usual probability properties](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent)
* 2026-08-06 06:16:52Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent/comments/kGdvgrn2idk9Su6vc](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent/comments/kGdvgrn2idk9Su6vc)
* Comment URL (HTML): [/posts/GR3TqJebMsiLQBqjv/ii-anthropic-reasoning-with-duplication-is-not-consistent/comment/kGdvgrn2idk9Su6vc](/posts/GR3TqJebMsiLQBqjv/ii-anthropic-reasoning-with-duplication-is-not-consistent/comment/kGdvgrn2idk9Su6vc)
I agree about 75% with this perspective.
### Comment by [Stuart_Armstrong](/users/stuart_armstrong) on [II. Anthropic reasoning with duplication is not consistent with the usual probability properties](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent)
* 2026-08-05 16:21:38Z
* Karma: 4
* Total votes: 2
* Comment URL (Markdown): [/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent/comments/BZpg22BtcEZZ9YAX2](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent/comments/BZpg22BtcEZZ9YAX2)
* Comment URL (HTML): [/posts/GR3TqJebMsiLQBqjv/ii-anthropic-reasoning-with-duplication-is-not-consistent/comment/BZpg22BtcEZZ9YAX2](/posts/GR3TqJebMsiLQBqjv/ii-anthropic-reasoning-with-duplication-is-not-consistent/comment/BZpg22BtcEZZ9YAX2)
Ok, got the issue. Thanks!
The condition I'm calling "simple Bayes" is that, if there is no doubt as to which agent you are in the world, then proceed by taking the prior over worlds and updating on the evidence "there exists a person in this world who have made this observation" (in non-anthropic situations, "there exists a person who has observed history H" and "I have observed history H" contain the same information).
On Monday, there is uncertainty are to which agent SB is. On Tuesday there is not: the two agents in the Tails world can tell each other apart, based on room number observation.
So simple Bayes applies to Sunday and Tuesday, but not to Monday.
On Sunday, the observations are independent of the coin flip, so P("I saw Sunday" | w_T) = P("I saw Sunday" | w_H) = 1.
If h_0 = ("I saw Sunday"), then P(w_T | h_0) = P(h_0 | w_T)P(w_T)/((P(h_0 | w_T)P(w_T) + P(h_0 | w_H)P(w_H)) = 1*(1/2)/(1/2+1/2)=1/2. Same thing for P(w_H | h_0).
On Tuesday, in the tails world, there exists, with certainty, an agent who has observed: h_1=("it's Sunday", "I'm awake on Monday", "it's Tuesday and my room number is 1"). Similarly, in the heads world, there exists, with certainty, an agent (the only agent) with the same observation.
Since both exist with certainty, then the formula is the same as before: P(w_T | h_1) = P(h_1 | w_T)P(w_T)/((P(h_1 | w_T)P(w_T) + P(h_1 | w_H)P(w_H)) = 1*(1/2)/(1/2+1/2)=1/2, and the same for P(w_H | h_1).
So the naive direct application of Bayes in non-anthropic situations gives us these values. The impossibility result is just that, given these values for Sunday and Tuesday, there are no values on Monday (the time-slice where we can't use simple Bayes because there is genuine uncertainty as to which agent SB is) that allow the martingale condition to extend between Sunday and Tuesday.
Note that I'm not saying that SIA or SSA are wrong or that you can't do anthropic probability. I'm saying that if you do do anthropic probability, you have to drop some intuitive properties possessed by standard probability.
For instance SIA drops the martingale condition from Sunday to Monday (indeed SIA always obeys simple Bayes). SSA (which is less uniquely defined) either drops the martingale condition from Monday to Tuesday or drops simple Bayes on Tuesday (using "centered worlds" is an explicit acknowledgement of dropping simple Bayes).
### Comment by [Stuart_Armstrong](/users/stuart_armstrong) on [II. Anthropic reasoning with duplication is not consistent with the usual probability properties](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent)
* 2026-08-05 14:09:15Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent/comments/r5f8L2aijnD72KYDy](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent/comments/r5f8L2aijnD72KYDy)
* Comment URL (HTML): [/posts/GR3TqJebMsiLQBqjv/ii-anthropic-reasoning-with-duplication-is-not-consistent/comment/r5f8L2aijnD72KYDy](/posts/GR3TqJebMsiLQBqjv/ii-anthropic-reasoning-with-duplication-is-not-consistent/comment/r5f8L2aijnD72KYDy)
To get the values of Q(T| awake, room 2) and Q(T| awake, room 1), I am applying "simple Bayes" because upon seeing awake and room 1, there is no longer any anthropic uncertainty as to who you are in the world.
Then, formally, I am assuming that Q is a well defined probability in general (obeying all the Bayesian equations, amongst others) and aiming to get a contradiction from this assumption.
So I set up an equation involving terms like Q(room 1 | awake) in order to infer what these would have to be. I've used simple Bayes to pin down some of the terms, and then I'm using the assumption that Q is well defined to set up an equation between the known terms and the unknown terms.
So one can write:
Q(T| awake, room 1) = (1/2) (t) / (1-t/2) = t/(2 - t)
And then, plugging in the known value of Q(T | awake, room 1), one deduces Q(T | awake) = 2/3. I haven't assumed that value; I've deduced it, under the "Q is reasonable, martingale, and simple Bayes works" assumptions (actually, all that I can deduce - and all that I need - is that Q(T | awake) > 1/2; getting it to be 2/3 requires *slightly* stronger symmetry assumptions).
I can also deduce Q(T | awake) = 1/2 (martingale from Sunday to Monday), and thus get a contradiction. One of my assumptions must be wrong. Q cannot simultaneously obey Bayesian updating, obey the martingale, and obey simple Bayes.
### Comment by [Stuart_Armstrong](/users/stuart_armstrong) on [III. Anthropic reasoning has issues with infinite worlds; D-SIA can fix this](/api/post/iii-anthropic-reasoning-has-issues-with-infinite-worlds-d)
* 2026-08-05 13:58:51Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/iii-anthropic-reasoning-has-issues-with-infinite-worlds-d/comments/hmL4SsDseMZggdAGF](/api/post/iii-anthropic-reasoning-has-issues-with-infinite-worlds-d/comments/hmL4SsDseMZggdAGF)
* Comment URL (HTML): [/posts/wZFxgrQmA8Gfmw7AG/iii-anthropic-reasoning-has-issues-with-infinite-worlds-d/comment/hmL4SsDseMZggdAGF](/posts/wZFxgrQmA8Gfmw7AG/iii-anthropic-reasoning-has-issues-with-infinite-worlds-d/comment/hmL4SsDseMZggdAGF)
Yep.
I'm wondering if the finite speed of light + planks constant giving a minimum coarsness to the universe + the cosmological constant giving a maximal size to the universe + the landauer limit preventing infinite computations in a finite universe... might not be strong evidence for a non-infinite universe?
### Comment by [Stuart_Armstrong](/users/stuart_armstrong) on [II. Anthropic reasoning with duplication is not consistent with the usual probability properties](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent)
* 2026-08-05 13:52:29Z
* Karma: 4
* Total votes: 2
* Comment URL (Markdown): [/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent/comments/orbJ4gzwFgHiwej2u](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent/comments/orbJ4gzwFgHiwej2u)
* Comment URL (HTML): [/posts/GR3TqJebMsiLQBqjv/ii-anthropic-reasoning-with-duplication-is-not-consistent/comment/orbJ4gzwFgHiwej2u](/posts/GR3TqJebMsiLQBqjv/ii-anthropic-reasoning-with-duplication-is-not-consistent/comment/orbJ4gzwFgHiwej2u)
Yep, the martingale is satisfied for both.
It's the equivalence between (conditional duplication) and (unconditional duplication + conditional killing) that I'm disputing (or at least saying that you can't just assume for free).
The difference is clear for the formal martingale: let's count the full agent histories from beginning to end. Conditional duplication has three histories - (Sunday, awake, room 1, T), (Sunday, awake, room 2, T) and (Sunday, awake, room 1, H). While (unconditional duplication + conditional killing) has four - those three plus extra one heads: the "death" history of just (Sunday).
Thus the martingale in the (ud + ck) case has an extra possible future observation to consider: specifically, the non-observation. That's why "awake" gives information in the (ud+ck) case (not all your duplicates would be guaranteed to see it) but not in the (cd) case (where all your duplicates are guaranteed to see it).
(It doesn't help if one says that "non-observations don't count", because then the theory breaks breaks the martingale in the case of just "conditional killing" on its own.)
Now, philosophically, I'm partial to the idea that "instant duplication followed by deletion" should be the same as "not creating duplicates at all"; that's a strong pro-SIA argument. But that doesn't fix the martingale; it instead shows that *if we assumed* that there were instantly-deleted duplicates, then that assumption would fix the martingale.
### Comment by [Stuart_Armstrong](/users/stuart_armstrong) on [Sleeping Beauty as a Mind Killer](/api/post/sleeping-beauty-as-a-mind-killer)
* 2026-08-04 17:09:02Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/sleeping-beauty-as-a-mind-killer/comments/rYuYqHkPuHbK6s4Mo](/api/post/sleeping-beauty-as-a-mind-killer/comments/rYuYqHkPuHbK6s4Mo)
* Comment URL (HTML): [/posts/6tKs9DPzNdBRhHJ8T/sleeping-beauty-as-a-mind-killer/comment/rYuYqHkPuHbK6s4Mo](/posts/6tKs9DPzNdBRhHJ8T/sleeping-beauty-as-a-mind-killer/comment/rYuYqHkPuHbK6s4Mo)
But what is the mechanism by which the doomsday argument works? The Laplace example has a clear prior powering it, but why should it apply here?
### Comment by [Stuart_Armstrong](/users/stuart_armstrong) on [I. Anthropic reasoning without duplicates is just standard Bayesian updating](/api/post/i-anthropic-reasoning-without-duplicates-is-just-standard)
* 2026-08-04 15:07:58Z
* Karma: 4
* Total votes: 2
* Comment URL (Markdown): [/api/post/i-anthropic-reasoning-without-duplicates-is-just-standard/comments/ZCYvNsBFmAaZrdSwq](/api/post/i-anthropic-reasoning-without-duplicates-is-just-standard/comments/ZCYvNsBFmAaZrdSwq)
* Comment URL (HTML): [/posts/wYgZnEQicfCGuEEC7/i-anthropic-reasoning-without-duplicates-is-just-standard/comment/ZCYvNsBFmAaZrdSwq](/posts/wYgZnEQicfCGuEEC7/i-anthropic-reasoning-without-duplicates-is-just-standard/comment/ZCYvNsBFmAaZrdSwq)
The alternative I found, D-SIA, does extend to infinite universes but isn't strongly fond of them (it likes large worlds, but doesn't have an infinite update towards infinite worlds).
### Comment by [Stuart_Armstrong](/users/stuart_armstrong) on [II. Anthropic reasoning with duplication is not consistent with the usual probability properties](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent)
* 2026-08-04 15:06:34Z
* Karma: 2
* Total votes: 1
* Comment URL (Markdown): [/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent/comments/GtDBSLw6ykdRJAJDx](/api/post/ii-anthropic-reasoning-with-duplication-is-not-consistent/comments/GtDBSLw6ykdRJAJDx)
* Comment URL (HTML): [/posts/GR3TqJebMsiLQBqjv/ii-anthropic-reasoning-with-duplication-is-not-consistent/comment/GtDBSLw6ykdRJAJDx](/posts/GR3TqJebMsiLQBqjv/ii-anthropic-reasoning-with-duplication-is-not-consistent/comment/GtDBSLw6ykdRJAJDx)
Yep, roughly where I'm at.
### Navigation
* [Front page](https://www.lesswrong.com/api/home)
* [Markdown API documentation](https://www.lesswrong.com/api/SKILL.md)