I think this can't be right, because counting cannot substitute for probabilities even in simple cases. Flip a biased coin. Now there are two worlds, but their probabilities aren't equal, and might not even have a simple ratio like 70/113 or whatever. This can be done with very basic physics, e.g. initialize a qubit, rotate it by some angle and measure it.
You can't just do sums with hyperreals, you can do integrals too. Maybe I should've emphasized that.
A quantum multiverse is still one mathematical structure, it's just one element of Tegmark IV. By talking about adding up all "worlds", I was proposing to add up all universes (roughly speaking) - obviously within a universe you might wanna do integrals instead. Of course you can't just arbitrarily carve the quantum multiverse in two differently-sized parts, call each one a "world" and say they both have the same probability, that would be absurd and I'm not proposing that.
I didn't mean to address the question of what probabilities are in full generality, just on the level of abstraction of metaphysics / Tegmark IV / UDASSA / etc. I guess for a quantum coin in our universe (which is a more "zoomed in" context), I don't really have anything new to say - just whatever quantum mechanics does, which is an integral over a measure in some way, I think? And then you derive your actual betting odds from your preferences if necessary, e.g. in Sleeping Beauty. (so I guess it's kind of a combination of a reality fluid and a caring measure).
Ok, I think I understand it a bit better now. But one thing I'd like to note is that UDASSA wasn't "zoomed out". One cool feature of the UD is that, given enough observations from our universe (but not the laws of physics), it'll reconstruct the right distribution for more observations from our universe. It doesn't need separate levels for "universes" and "stuff within a universe", it just does everything on one level. So if you replace it with a two-level system, that might be a bit unsatisfying.
Cool, thank you for putting in the effort!
Yeah, that's a good point - I should clarify that. Technically, I'm not summing over universes either (that's why I said "roughly speaking" earlier), but over all possible computations that lead to my observations, just like UDASSA.
The crux in whether this reconstructs reasonable actions / betting odds (e.g. that it generally converges with quantum mechanics) too is what I mentioned in the post - whether it reconstructs some version of a simplicity prior.
Speculatively, we might even recover some version of a simplicity prior from it, since simple worlds might reoccur more frequently across all possible computations, i.e. influence the infinite sum more.
You can get the Solomonoff simplicity prior just by taking a uniform prior over programs of length L on a plain UTM and letting L tend to infinity. See result 3.8.1 in Hutter's An Introduction to Universal Artificial Intelligence.
(I don't think you need hyperreal numbers to prove this result.)
You can get the Solomonoff simplicity prior just by taking a uniform prior over infinitely long programs.
How is that statement different from the statement I made?
It doesn't privilege length by taking any explicit limit in length at all.
How is 'uniform prior over programs of infinite length' defined if not via a limit in length?
You sample a program of infinite length (by, for each index, sampling a random instruction to go at that index), and then you run it. Of course, with probability 1 only finitely many of the instructions will be read, since there are infinitely many chances to encounter some trap for the instruction pointer.
Nice. Is there a proof for that written up somewhere public?
No, most of my thinking about that was today. I'd be interested in a programming language where infinitely long programs have probability <1 of being equivalent to some finite program and have dynamics other than "keep running random instructions, achieve nothing of consequence" like what happens when you sample the target address for where your goto jumps to uniformly at random.
Potentially the structure/topology of address space matters in the infinite-program setting? That way may lie a derivation of physics.
I think any natural-number-indexed language where a nonzero fraction of instructions is "jump back 100 instructions" has almost all infinite programs equivalent to some finite program.
Just tried to prove it and I think you're right. For example let's say the instructions are "jump forward 1337" and "jump back 100". Then, no matter where you are, "jump forward 100 times and then back 1337 times" would lead to a loop. So there's a small but fixed chance that the immediately next instructions define exactly this loop, unless we run into a previously visited instruction which means we're in a loop anyway. So over infinite time, almost all trajectories will fall into this loop or some other one. The proof generalizes to all languages in the obvious way.
EDIT: It's fun to think about the limits of applicability of this. All that is needed is that instructions are i.i.d., all jumps are relative, and the chance of a loop is nonzero. Other than that we're pretty free. For example, we could have countably many instructions: jump forward by BB(1) with probability 1/2, jump back by BB(2) with probability 1/4, jump forward by BB(3) with probability 1/8, jump back by BB(4) with probability 1/16... Then the expected time to loop is finite and not too large, but the expected distance to the loop is an infinity beyond comprehension.
This is a vague connection and possibly a misunderstanding, but the idea of imagining everything being inside one program and deriving physics from that kind of sounds like the Universal Dovetailer Argument, in case you haven't heard of it.
That sounds applicable to any discussion of the Solomonoff prior at all; what I meant by deriving physics is that, by analyzing Turing-complete machine models for e.g. how well they handle infinities we might privilege hypotheses like our spacetime enough for papers like Max Tegmark's On the dimensionality of spacetime to take us the rest of the way.
Hm, I want to react to this with "Difficult to Parse", but I notice this feels a bit too accusatory. I don't know if it's objectively difficult to parse or if I just lack knowledge in this area. I'd like there to be a reaction like "I don't understand" that's only about my own experience.
I haven't encountered this result, but it makes intuitive sense to me that something of this form could work to define the prior too.
Okay, talked to Fable for a bit and understand it now I think, cool! Thanks Gurkenglas.
Cool result, thanks. But this already smuggles in description length dependence (i.e. something in the direction of a simplicity prior) by requiring the prior to be uniform over programs of each length, no?
(This is a general confusion for me when it comes to the "we can derive Occam's razor from nothing" slogan)
In this formulation, you only consider programs of exactly length
One could still object that we are privileging length by taking any explicit limit in length at all. But I dunno, this seems pretty practically motivated to me.
EDIT: Sol says yes,
Sol:
Yes — it converges to the same Solomonoff semimeasure. In the book’s notation, let
, so Section 3.8.1 proves . If a program is sampled uniformly from all bitstrings of length at most , then .
Sinceis increasing, . For any fixed , .
Thus, while for every , hence . Therefore . The same conclusion also holds for the other possible interpretation—choose a length uniformly from , then a string uniformly at that length—by the ordinary Cesàro convergence theorem. The relevant source is pp. 159–161 of the authors’ PDF: https://www.hutter1.net/publ/uaibook2.pdf. One caveat: this uses exactly the book’s model, where the candidates are all binary strings and
ignores unread padding up to the point is printed. A restriction to an arbitrary set of “syntactically valid programs” would need separate assumptions on how their counts grow.
Cool! Thanks for checking that. Made me talk to Fable for a bit and understand it better too.
One could still object that we are privileging length by taking any explicit limit in length at all. But I dunno, this seems pretty practically motivated to me.
I would probably still have this objection, yeah (especially now that I understand better how it results in the simplicity prior, with shorter programs having exponentially more ways to be padded) - we are at a level of fundamental philosophy where IMO we are looking for theoretical principledness, not pragmatic appeal. But I agree that I probably undersold the UD a bit in the post. It's quite natural, and there's definitely a deep insight there about how natural a simplicity prior is.
(EDIT: Gurkenglas makes a good counterargument!)
Sounds fun! I suspect you cannot get a simplicity prior back out, but happy to follow along.
The sensitivity of infinite sums to ordering seems to me to be fatal for the idea of attaching values to arbitrary infinite sums. People have tried in the context of attaching infinite utilities to paradoxical games like St. Petersburg, but they have never got very far. There are theorems about the impossibility of defining well-behaved preference relations over the set of all probability distributions over some space; for example.
Ord discusses St. Petersburg in the paper!
I think this is not as bad as it first seems - I will try to write my next post on this.
I think objection one and two could be related some non-trivial way. Essentially, can we encode conditional probabilities with these order sensitive sums. This then implies complexity which makes Boltzmann brains unlikely, although this will not be simple in general.
Second, I would really be interested in building some good intuition for what the order commutators tell us ie do they have a measure?
This then implies complexity which makes Boltzmann brains unlikely, although this will not be simple in general.
Interesting, how so? I don't see the reasoning yet.
I would really be interested in building some good intuition for what the order commutators tell
Check out my post about Ord's paper (I link a Claude chat there) and the original paper itself!
This post is somewhat niche, and I will sometimes not give context or link relevant background.
There’s a big debate that has played out in slow motion on LessWrong over the past two decades, between two broad ways of putting a measure over all possible realities (Tegmark IV):
These both have significant drawbacks:
Unfortunately, there are infinite possible worlds and every event happens infinitely many times - so we do need some kind of measure to calculate probabilities and the effects of our actions.
Or do we?
Recently, I came across Toby Ord’s “Evaluating the infinite” paper from last year. I wrote about my reaction to it here, and here’s Ord’s Twitter summary - the gist of it is that using hyperreal numbers (where infinity + 1 does not equal infinity) to evaluate infinities in various fields is actually more promising and coherent than people previously thought.
I think this paper might have gone under the radar a bit. I couldn’t find any discussion of it on LessWrong, for example.
More than the specific hyperreal formalism, my main takeaway was more philosophical - a sort of “scales falling from my eyes” / “paradigm shift” realization that the ontology of “infinity + 1 = infinity” never really made sense in the first place. I can’t even remember why I believed it for so long, like it was just an unexamined assumption that immediately collapsed when I thought about it for a moment.
Here’s a quote from an email exchange with Ord that I found useful (reproducing it with permission):
It does feel to me now like the default way of thinking should be that infinity + 1 > infinity. Like… obviously if you add 1 to something it becomes bigger?
Let’s take this back to the measure problem. If we take this semi-philosophical stance seriously, why do we even need a measure?
Could we just sum up all the infinite possible occurrences of our possible next inputs, normalize, and get probabilities about what our next input will be that way? Sum up everything that we care about across the infinite possible realities, and get estimates of the effects of our actions that way?
That sounds kind of insane. But the more I think about it, the more it feels like the only principled approach. It’s weirdly very grounded. Like, literally just add up everything? Details TBD?
That’s really all I wanted to get across in this post - that this approach seems to be very neglected among thinkers in this area. It very much doesn’t obviously fail, and nobody seems to have seriously thought about it or tried to work out its implications.[3]
It’s far more complicated and ambitious than the simple settings where Ord rigorously showed hyperreals work, and there’s other possible number systems where infinity + 1 > infinity (like surreal numbers), so I want to distinguish this from his specific hyperreal formalism. Let’s call it the measureless approach, for lack of a better term.
Speculatively, we might even recover some version of a simplicity prior from it, since simple worlds might reoccur more frequently across all possible computations, i.e. influence the infinite sum more.[4]
That said, it also has some serious issues (although for me personally, not enough to outweigh its appeal). For example:
So it’s definitely plausible that it will turn out to not make sense. But the existing approaches don’t seem clearly better!
So the measureless approach seems underrated to me.
The other major alternative is Schmidhuber’s speed prior.
They will to some extent, of course - see the ADT paper and Wei Dai here, or David Matolcsi’s Probabilities are not the right concept for the most exhaustive treatment that I know of.
Perhaps because of the quasi-ideology of cardinals that Ord talks about.
(arguably, this is already what the Solomonoff prior is doing, after all)
I think this is not as bad as it first seems - I will try to write my next post on this.
I think this is more natural than it first seems - see the 2nd edit in my post on Ord’s paper.