Stop doing decision theory without metaphysics

[Epistemic status: rant]

There’s something that annoys me about the reoccurring debates on decision theory in this corner of the internet.

Take a simple blackmail scenario:

Omega, a near-perfect predictor, knows a piece of embarrassing information about you. He threatens you that he will release it to the public if you don’t pay him 100$. However, making the threat is slightly costly to him, and he wouldn’t have done it if he hadn’t predicted that you would pay. Do you pay?

Let’s say we want to argue for the Functional Decision Theory (FDT) answer that you shouldn’t pay. It was originally motivated by the observation that such agents seem to achieve higher utility (“rationality is about winning”), since they don’t get blackmailed in the first place. But that only leads to making yourself into such an agent in advance (which everybody generally agrees you should do[1]) - it doesn’t clearly apply when you are already being blackmailed and have never thought about the question before, or if you are an AI who was just created and is instantly blackmailed before being able to self-modify or make precommitments.

I see three broad ways to make FDT’s recommendation make sense in that case[2]:

  • Alternate realities: You are in a branch of reality where you got blackmailed, but other branches exist, and you can make it so you don’t get blackmailed there. For example, this is how people usually talk about Updateless Decision Theory (UDT), and Paul Christiano seems to understand it like this[3].

  • Changing the past /​ making this reality impossible: There is only one reality, but deciding not to pay retroactively makes it so you didn’t get blackmailed in the first place /​ makes the reality where you get blackmailed impossible, and makes it so something else actually happens. For example, Joe Carlsmith here at least explores this direction, and there’s Timeless Decision Theory (TDT)-style framings like Nate Soares here and here.

  • Platonism: You might be in the abstract object of the computation that is your decision process, which determines whether you get blackmailed in actual reality. For example, this is arguably the most natural way to think about the FDT paper’s causal graphs, where the node with your decision function is upstream of the physical world.[4]

These are all metaphysical claims—therefore, you need metaphysics to make FDT make sense.

(I understand many people will disagree with me on this—I think it’s a widespread misconception. I address some counterarguments in this footnote[5])

And I want to be absolutely clear here—I am very sympathetic to these claims! In my personal opinion, something in the vicinity is actually true. But… they are metaphysical claims, and ones that most people probably don’t immediately accept.

So let’s say that someone only knows this fact about the topic, and and starts reading what people are saying about it online. Imagine their surprise at finding that the back-and-forths between FDTers and anti-FDTers are largely not about metaphysics. Huh?

What the hell is going on?

I understand that metaphysics is really annoying and hard to think about, but… guys, I think this is your actual crux!


Take the comment section under Bentham’s Bulldog’s recent post about FDT. There are such luminaries as Scott Alexander and Stuart Armstrong chiming in. Yet barely anyone is bringing up metaphysics—instead the discussions circle around semantic disagreements about what the words “rational” and “decision theory” should mean, confusion that FDT is necessary for everyday psychological precommitments[6], confusion that everyday psychological precommitments are sufficient to become an FDT agent[7], confusion that the observation that FDT agents get more utility is enough to fully justify FDT[8], and so on and so forth.

No! What are you doing? You have one very concrete disagreement—talk about that one! There is only one spot where FDT and a normal worldview diverge, and it’s the kind of scenario above, where you have no time to self-modify/​precommit. For everything else, normal decision theory is sufficient![9]

And if you talk about that kind of scenario, you’ll be able to actually get to the bottom of your disagreements—which is metaphysics.

Concretely, my message to FDT proponents is this: Start being clear to yourself, and to others, about your metaphysical stances. You are confusing everyone by leaving them implicit.

I have a suspicion that a substantial amount of the resistance to FDT is that people can tell that you’re doing something fucky. They can tell that FDT doesn’t really make sense without additional metaphysical commitments, and that you’re just pretending it does. But if you were to say “you might just be inside the abstract computation” or whatever, I think people would agree that FDT makes sense given that assumption. And then you can actually talk about whether those metaphysics are reasonable and how to handle that, instead of talking past each other.

Old-school LessWrong was more explicit about this stuff. There was plenty of open talk, back when UDT was the flagship decision theory, about crazy stuff like everything existing, that we can decide what gets to “exist”, supernatural voices from the sky, and so on. My guess is that a lot of those veterans from back then are aware of what I’m saying here, and have just gotten tired of talking about metaphysics and don’t chime in much anymore.

For example, Paul Christiano:

“What the physical world actually looks like” seems likely to be a matter of preferences… probably I’d go with one of “avoid thinking about it“ or “move away from ascribing metaphysical significance to existing,” or at least moving towards a state of ignorance.

(on his take on decision theory) I justify that perspective in significant part from a position of radical uncertainty: I’m not sure if I’m thinking about worlds that don’t exist, or if it’s us who don’t exist and there is some real world somewhere thinking about us.

Paul Christiano is not the type of writer to say something that wild if it’s not necessary for what he’s arguing for! It’s really not as simple as some of you think!

I think something has gotten lost somewhere, in the transition to the more narrow, academic framing of FDT. Some of the newer arrivals to the space didn’t get the message that there is an underlying metaphysical question, and are going around thinking that it’s a normal decision theory like any other. My past self from three years ago definitely didn’t fully realize what it was arguing for when it was arguing for FDT.

I mean, let’s just take a look at these apparent facts about the paper (emphasis mine):

Scott Garrabrant: I think that there is this (obvious to LessWrongers, because it is deeply entangled with the entire LessWrong philosophy) ontology in which “I am an algorithm” rather than “I am a physical object.” I think that most decision theorists haven’t really considered this ontology. I mostly view FDT (the paper) as a not-fully-formal attempt to bridge that inferential difference and argue for identifying with your algorithm… the heart of FDT is about the algorithm question.

Rob Bensinger: Agreed the FDT paper was mainly about the algorithm axis.

The FDT paper is close to my heart and I respect the authors a lot, but… if this was the intention, why are you writing a pure decision theory paper? That’s a metaphysical claim. And there is no mention of metaphysical commitments in the paper at all—in fact, in the conclusion, it explicitly claims that it all works without any metaphysics (which just isn’t true, in my opinion—maybe on a formal level, but not on a philosophical level). That seems almost deceptive—no wonder that this confuses everyone, including the referees of the paper.

Wei Dai (not claiming that he would fully agree with my take here):

I feel like MIRI perhaps mispositioned FDT (their variant of UDT) as a clear advancement in decision theory, whereas maybe they could have attracted more attention/​interest from academic philosophy if the framing was instead that the UDT line of thinking shows that decision theory is just more deeply puzzling than anyone had previously realized.

Yeah—I could imagine another version of the paper, something like “computationalism applied to decision theory leads to weird metaphysical challenges like seemingly changing the past and difficult formal problems like logical counterfactuals if we take it seriously. But we should—because it’s very elegant and solves a lot of issues (e.g. dynamic inconsistency, coordination with similar agents), and computationalism is a widespread position.”[10]

At this point, I see the failure of academia to come up with FDT before LessWrong partially as the same academic failure mode, of too little interdisciplinarity, that caused the FDT paper authors to feel pressure to stay unnaturally agnostic about metaphysics. In retrospect, it’s obvious that the concepts of logical causation and updatelessness have been grasped at in academia for a long time[11], but that they just never fully got there—because it needs metaphysics to actually make sense. But decision theory and metaphysics are different subfields of philosophy, so bringing them together doesn’t come naturally. Grand unifying theories are institutionally difficult, and often actively disincentivized. LessWrongers, on the other hand, were not academics, and nonconformist enough to put together the obvious pieces lying around[12].

But then, let’s be clear about what we’re actually doing! Let’s be proud of the metaphysics again! Let’s be proud of being generalist interdisciplinary thinkers!

Getting to this point via the heuristic of “rationality is about winning” was perfectly valid. We did a great job! We outdid the academics! But that just means we get to argue about metaphysics now[13]. We leveled up.


Finally, I’d like to end on a positive note—Nate Soares in Notes on “Can you control the past” is who I’ve seen make things explicit the most in the last few years (although still not as much as I would like). For example, he says:

I agree “you’re flat-out metaphysically wrong (in a way that seems even worse than violating [the principle of not passing up] guaranteed payoffs)” is a valid counterargument to my actual position (in a way that “you violate [the principle of not passing up] guaranteed payoffs” is not). :-)

Then Joe Carlsmith doesn’t press him any further on this, presumably in part, again, because metaphysics seems too annoying to get into. But… yes! More of this please!

Thanks to Hein de Haan and @eigengender for valuable discussion of these ideas.

  1. ^

    even by academic decision theorists—although they might not fully realize the implications, e.g. that you need to become the kind of person who would choose to burn to death in Bomb*.

  2. ^

    Assuming the predictor isn’t predicting you by simulating you with high fidelity (since otherwise you can just say that you might be in the simulation). This is reasonable to postulate because FDT is also supposed to change your action with very bad prediction on the part of Omega, e.g. with only a 60% success rate.

  3. ^

    “I’m not sure if I’m thinking about worlds that don’t exist, or if it’s us who don’t exist and there is some real world somewhere thinking about us.”

    Note that it’s also technically possible to justify this in a purely axiological way, that you just also care about alternate realities, without making a metaphysical claim as to whether they’re real. But I don’t buy this, this is just a formal trick—you wouldn’t care about something that you don’t, on a gut level, believe is “real” in some way.

  4. ^

    It’s a little ambiguous, but from the paper: “What’s remarkable about this line of reasoning is that even in the case where Fiona has observed that box B is full, when she envisions two-boxing, she envisions a scenario where she instead (with high probability) sees that the box is empty. In words, she reasons: “The thoughts I’m currently thinking are the decision procedure that I run upon seeing a full box. This procedure is being predicted by the predictor, and (maybe) implemented by my body. If it outputs onebox, the box is likely full and my brain implements this procedure so I take one box. If instead it outputs twobox, the box is likely empty and my brain does not implement this procedure (because I will be shown an empty box). Thus, if this procedure outputs onebox then I’m likely to keep $1,000,000; whereas if it outputs twobox I’m likely to get only $1,000. Outputting onebox leads to better outcomes, so this decision procedure hereby outputs onebox.”

  5. ^

    (This footnote is very long, so probably actually click on it instead of hovering over it). Three out of three non-experts that I spoke to about this had this misconception. It makes me think that the metaphysics-agnostic framing in the FDT paper might genuinely have had bad effects here. Anyway, here are some counterarguments and alternative framings I’ve encountered, and my responses:

    • More on impossible realities: “But standard academic decision theories need to consider counterpossibles too, since every counterfactual presupposes something that isn’t true (that we might do something other than what we actually will). So there’s no more metaphysics than in a normal worldview”. Okay, sure, but they don’t imagine that they can make their own reality impossible—only FDT does that (academic decision theories only imagine inconsistencies insofar as they can actually make them consistent after all, by choosing that option). That, to me, is a distinct stance—a much more metaphysical one. It doesn’t make sense in a normal worldview.

      Then the response is often “but the impossible realities never actually happen, it’s purely concentrated in counterfactuals”. Well, sure, but the consideration is always present in this framing. If Omega is imperfect, you’re making decisions by imagining an x% chance that you’re making this reality impossible with your action (and that something else actually happened). What you’re really saying (like Nate Soares here, ctrl-F “metaphysically”) is “You need to consider crazy metaphysics (imagining you’re in a hypothetical impossible reality) to make decisions in a reasonable way”, and like… yes, exactly. That’s de facto a metaphysical stance. “It makes things easier if we imagine this” is not a decisive argument for a metaphysical claim that permits you to leave it unstated as if it’s obvious. (This line of argument sometimes feels like people are really trying their absolute hardest to pretend they’re not taking unusual metaphysical stances, even though they are.)

    • Pseudo-simulation (this is the other way I see to interpret the causal graph framing in the FDT paper): “Even if Omega isn’t predicting you by simulating you with high fidelity, if your decision function is simple and doesn’t take your full sense experience into account, the computation that you actually are (your decision function) would by definition still exist in Omega’s head, even if your full self doesn’t exist—so there’s no paradox since the computation that you actually are (your decision function) is still instantiated somewhere.” Nate Soares frames it like this here, and I’ve encountered it in personal conversation. This is a useful framing in general, but I don’t think it ends up working for this.

      If we assume I’m actually being blackmailed (like the problem statement says), I have access to my full sense experience. So naively, I have very high credence in the belief that my sense experience exists /​ is actually instantiated somewhere. But the reasoning above is predicated on thinking that it might not be. In fact, to make the math work out the same way, you have to be completely agnostic about whether your sense experience is happening, as if you have no information about it whatsoever (since even a little bit of updating on it would bias you towards paying more than FDT recommends). To be clear, not agnostic about whether your sense experience is accurate about the external world—agnostic about whether it’s happening at all. That’s not technically metaphysical, I suppose, but it’s a very radical epistemological stance that most people would still disagree with. (it rejects Descartes!) Also, it’s just obviously wrong, and you don’t actually believe this. The people I’ve talked to have all also denied that this kind of radical skepticism is needed to justify FDT—I’m including it here more for completeness.

      (Note that I am not saying here that you can fool Omega by taking your full sense experience into account in your decision—I’m purely analyzing what this reasoning actually looks like, and whether it makes sense as stated, or whether there is an additional claim needed to reach the correct answer that you shouldn’t pay.)

      Again, then you might say “but this never actually happens”, and, no, even when Omega is imperfect, in this framing you have an x% credence that your sense experience isn’t actually happening.

    A few other, less common counterarguments:

    • “By posing this problem, you are already assuming that you are in a situation that is vanishingly unlikely by FDT’s lights—so it’s not really a good counterargument that it performs badly there by not paying.”: I think that this is a good intuition aid, but it’s not really an explanation that would make sense to the person already being blackmailed. It doesn’t clearly matter to them that the situation they are in is “unlikely” in some abstract sense. For that, you need one of the metaphysical claims.

    • “We need to deflate/​dissolve the notions of existence and instantiation.” My personally preferred solution, but also clearly a metaphysical move.

  6. ^

    No—you can decide to e.g. become a virtuous, reliable agent on entirely causal reasoning.

  7. ^

    No—it only leads to son-of-CDT (which, to be fair, is pretty powerful, as I discuss here, but fails on the blackmail scenario above).

  8. ^

    It’s not, as we see in the blackmail scenario above.

  9. ^

    Note for example, that XOR Blackmail—in the way it’s usually understood, which is generally taken to disprove EDT—is such a problem. So it’s really more of a counterexample to updatefulness, than to EDT-style counterfactuals.

    I am generally using FDT to mean updatelessness here—sorry for being imprecise (although FDT is supposed to be “an umbrella-term for UDT-ish approaches to decision theory”, so it’s pretty close. It was originally meant to be agnostic between EDT- and CDT-style counterfactuals, so I usually take the disagreement about those to not be central. And it probably doesn’t matter anyway, since “CDT=EDT?”.and updatelessness is the more important proposal, in my view.)

  10. ^

    (Joe Carlsmith hits some of these notes in Can you control the past?)

    But I can imagine other framings too.

  11. ^

    There’s shades of logical causation in the three papers here, and shades of updatelessness in Fisher’s and Gauthier’s disposition-based decision theories, Parfit’s “rational irrationality”, McClennen’s “resolute choice” (“it is rational to follow through on plans even when they become locally harmful”), Bratman (Intentions, Plans, and Practical Reason, 1987), and especially Meacham’s cohesive decision theory (“do as you would have bound yourself to do”) (footnote 34 is incredibly fascinating and prescient as an early discussion of the idea of updatelessness, and of how updateless to be).

  12. ^

    Of decision theory, metaphysics (e.g. Tegmark IV), and computationalism in philosophy of mind.

  13. ^

    Or talk about the very mundane commitment theory instead, my metaphysically minimalist version of LessWrong-style decision theory—your choice. But don’t do this weird in-between thing of arguing for fancy FDT and simultaneously pretending there’s no metaphysics going on. You can’t have your cake and eat it too.