PSA: There’s a third option in the “measure problem”
This post is somewhat niche, and I will sometimes not give context or link relevant background.
There’s a big debate that has played out in slow motion on LessWrong over the past two decades, between two broad ways of putting a measure over all possible realities (Tegmark IV):
Some “objective” prior (a “reality fluid”), usually a simplicity prior[1]: This is the position taken by Max Tegmark, Jürgen Schmidhuber and UDASSA.
A “caring measure”, where we say that our preferences determine our probabilities and maybe even what counts as “existing”. For example, Wei Dai here, Paul Christiano here and Scott Garrabrant.
These both have significant drawbacks:
A simplicity prior seems to imply some very counterintuitive things, like caring about people more the easier we can find them in the universe (and even weirder things, see David Matolcsi here and Joe Carlsmith here), and is partially dependent on an arbitrary choice of implementation (e.g. which Universal Turing Machine to use in UDASSA).
A caring measure just seems a bit unmotivated—intuitively, our probabilities (or existence itself) shouldn’t entirely depend on our preferences.[2] Ideally, we’d like something better.
Unfortunately, there are infinite possible worlds and every event happens infinitely many times—so we do need some kind of measure to calculate probabilities and the effects of our actions.
Or do we?
Recently, I came across Toby Ord’s “Evaluating the infinite” paper from last year. I wrote about my reaction to it here, and here’s Ord’s Twitter summary—the gist of it is that using hyperreal numbers (where infinity + 1 does not equal infinity) to evaluate infinities in various fields is actually more promising and coherent than people previously thought.
I think this paper might have gone under the radar a bit. I couldn’t find any discussion of it on LessWrong, for example.
More than the specific hyperreal formalism, my main takeaway was more philosophical—a sort of “scales falling from my eyes” / “paradigm shift” realization that the ontology of “infinity + 1 = infinity” never really made sense in the first place. I can’t even remember why I believed it for so long, like it was just an unexamined assumption that immediately collapsed when I thought about it for a moment.
Here’s a quote from an email exchange with Ord that I found useful (reproducing it with permission):
Part of the problem is that maths education has now instilled in many of us the counterintuitive principles of Hilbert’s hotel and the cardinal numbers, and that while these are a very useful concept of infinity, they are not the right one for this job and so our current mathematical intuitions are leading us more astray than if we were mathematically naive.
It does feel to me now like the default way of thinking should be that infinity + 1 > infinity. Like… obviously if you add 1 to something it becomes bigger?
Let’s take this back to the measure problem. If we take this semi-philosophical stance seriously, why do we even need a measure?
Could we just sum up all the infinite possible occurrences of our possible next inputs, normalize, and get probabilities about what our next input will be that way? Sum up everything that we care about across the infinite possible realities, and get estimates of the effects of our actions that way?
That sounds kind of insane. But the more I think about it, the more it feels like the only principled approach. It’s weirdly very grounded. Like, literally just add up everything? Details TBD?
That’s really all I wanted to get across in this post—that this approach seems to be very neglected among thinkers in this area. It very much doesn’t obviously fail, and nobody seems to have seriously thought about it or tried to work out its implications.[3]
It’s far more complicated and ambitious than the simple settings where Ord rigorously showed hyperreals work, and there’s other possible number systems where infinity + 1 > infinity (like surreal numbers), so I want to distinguish this from his specific hyperreal formalism. Let’s call it the measureless approach, for lack of a better term.
Speculatively, we might even recover some version of a simplicity prior from it, since simple worlds might reoccur more frequently across all possible computations, i.e. influence the infinite sum more.
That said, it also has some serious issues (although for me personally, not enough to outweigh its appeal). For example:
It doesn’t solve the puzzle of why we aren’t Boltzmann brains like UDASSA famously does.[4]
It seems like any formalism where infinity + 1 > infinity will give different results for an infinite sum depending on the order of the summands.[5] So this immediately leads to extreme ambiguity about which order to sum all possible realities in. (Although this might be just another way we recover some version of a simplicity prior, since plausibly the only principled order is one given by some UTM like in UDASSA—which would be a nice convergence).
But if some version of a simplicity prior is recovered, this might also recreate some of its issues.
So it’s definitely plausible that it will turn out to not make sense. But the existing approaches don’t seem clearly better!
So the measureless approach seems underrated to me.
- ^
The other major alternative is Schmidhuber’s speed prior.
- ^
They will to some extent, of course—see the ADT paper and Wei Dai here, or David Matolcsi’s Probabilities are not the right concept for the most exhaustive treatment that I know of.
- ^
Perhaps because of the quasi-ideology of cardinals that Ord talks about.
- ^
I think this is not as bad as it first seems—I will try to write my next post on this.
- ^
I think this is more natural than it first seems—see the 2nd edit in my post on Ord’s paper.
I think this can’t be right, because counting cannot substitute for probabilities even in simple cases. Flip a biased coin. Now there are two worlds, but their probabilities aren’t equal, and might not even have a simple ratio like 70⁄113 or whatever. This can be done with very basic physics, e.g. initialize a qubit, rotate it by some angle and measure it.
You can’t just do sums with hyperreals, you can do integrals too. Maybe I should’ve emphasized that.
A quantum multiverse is still one mathematical structure, it’s just one element of Tegmark IV. By talking about adding up all “worlds”, I was proposing to add up all universes (roughly speaking) - obviously within a universe you might wanna do integrals instead. Of course you can’t just arbitrarily carve the quantum multiverse in two differently-sized parts, call each one a “world” and say they both have the same probability, that would be absurd and I’m not proposing that.
Maybe I don’t understand your idea yet. You started with the question whether probabilities are a “reality fluid”, a “measure of caring”, or something else. How would you answer this question for the probabilities of a quantum coin in our universe?
I didn’t mean to address the question of what probabilities are in full generality, just on the level of abstraction of metaphysics / Tegmark IV / UDASSA / etc. I guess for a quantum coin in our universe (which is a more “zoomed in” context), I don’t really have anything new to say—just whatever quantum mechanics does, which is an integral over a measure in some way, I think? And then you derive your actual betting odds from your preferences if necessary, e.g. in Sleeping Beauty. (so I guess it’s kind of a combination of a reality fluid and a caring measure).
Ok, I think I understand it a bit better now. But one thing I’d like to note is that UDASSA wasn’t “zoomed out”. One cool feature of the UD is that, given enough observations from our universe (but not the laws of physics), it’ll reconstruct the right distribution for more observations from our universe. It doesn’t need separate levels for “universes” and “stuff within a universe”, it just does everything on one level. So if you replace it with a two-level system, that might be a bit unsatisfying.
Cool, thank you for putting in the effort!
Yeah, that’s a good point—I should clarify that. Technically, I’m not summing over universes either (that’s why I said “roughly speaking” earlier), but over all possible computations that lead to my observations, just like UDASSA.
The crux in whether this reconstructs reasonable actions / betting odds (e.g. that it generally converges with quantum mechanics) too is what I mentioned in the post—whether it reconstructs some version of a simplicity prior.
You can get the Solomonoff simplicity prior just by taking a uniform prior over programs of length L on a plain UTM and letting L tend to infinity. See result 3.8.1 in Hutter’s An Introduction to Universal Artificial Intelligence.
(I don’t think you need hyperreal numbers to prove this result.)
Cool result, thanks. But this already smuggles in description length dependence (i.e. something in the direction of a simplicity prior) by requiring the prior to be uniform over programs of each length, no?
(This is a general confusion for me when it comes to the “we can derive Occam’s razor from nothing” slogan)
In this formulation, you only consider programs of exactly length . I don’t know what happens if you instead take uniform prior over programs with length up to . My guess is that this would converge to the Solomonoff simplicity prior as well. The shorter programs are just too few in number to make much of a difference. I can ask Sol.
indeed converges to the Solomonoff semimeasure as well.
One could still object that we are privileging length by taking any explicit limit in length at all. But I dunno, this seems pretty practically motivated to me.
EDIT: Sol says yes,
Sol:
Cool! Thanks for checking that. Made me talk to Fable for a bit and understand it better too.
I would probably still have this objection, yeah (especially now that I understand better how it results in the simplicity prior, with shorter programs having exponentially more ways to be padded) - we are at a level of fundamental philosophy where IMO we are looking for theoretical principledness, not pragmatic appeal. But I agree that I probably undersold the UD a bit in the post. It’s quite natural, and there’s definitely a deep insight there about how natural a simplicity prior is.
Sounds fun! I suspect you cannot get a simplicity prior back out, but happy to follow along.
The sensitivity of infinite sums to ordering seems to me to be fatal for the idea of attaching values to arbitrary infinite sums. People have tried in the context of attaching infinite utilities to paradoxical games like St. Petersburg, but they have never got very far. There are theorems about the impossibility of defining well-behaved preference relations over the set of all probability distributions over some space; for example.
Ord discusses St. Petersburg in the paper!
I think objection one and two could be related some non-trivial way. Essentially, can we encode conditional probabilities with these order sensitive sums. This then implies complexity which makes Boltzmann brains unlikely, although this will not be simple in general.
Second, I would really be interested in building some good intuition for what the order commutators tell us ie do they have a measure?
Interesting, how so? I don’t see the reasoning yet.
Check out my post about Ord’s paper (I link a Claude chat there) and the original paper itself!