Reframing LessWrong-style decision theory as “commitment theory”
Thanks to @eigengender and especially Chris Lakin and Simon Dima for valuable comments on a draft.
In this post, I will present an alternate framing[1] of LessWrong-style decision theory. I believe it should make its concrete recommendations much more palatable to people who find them absurd. Very little of the theoretical substance here is original to me, and the core claim will be obvious to many, but I don’t think anyone has cleanly framed it all like this for easy consumption.
TLDR: Framing the thorny issues of decision theory[2] as stemming from a novel, non-decision-theoretic consideration (something like “cooperating with your past self”) - instead of forcing the relevant intuitions and claims into a decision theory frame—is much less counterintuitive, and gives exactly the same recommendations in practice.
Introduction
Let’s briefly check out Will MacAskill’s “Bomb*” scenario[3]:
You’re alone at home at the end of the universe. Omega—an amazing predictor of human behavior, but long since dead—gave you one last challenge. You face two boxes, Left and Right. You must take one of them.
In the Left box, Omega put either nothing or a live bomb; if there’s a live bomb, taking this box will set it off, setting you ablaze, and you will burn slowly to death. In the Right box, Omega put a glitter bomb; taking this box will set it off, covering you and your whole living room in glitter, and you’ll have to spend hours cleaning it up—very inconvenient.
Both boxes are transparent, so you can actually see the contents of the Left box—there is a live bomb.
Omega predicted whether you would choose Left or Right when faced with exactly this situation—of seeing a bomb in the Left box. If he predicted that you would choose Right, he put a bomb in Left (like you’re seeing now). If he had predicted that you would choose Left, he wouldn’t have put a bomb in Left, and the box would be empty.
You know that Omega has a failure rate of one in a trillion trillion in this situation.
You are the only person left in the universe. You have a happy life, but you know that you will never meet another agent again, nor face another situation where any of your actions will have been predicted by another agent. What box do you choose?
Functional decision theory (FDT), the main LessWrong-style decision theory, recommends taking the Left box—in the full knowledge that as a result you will slowly burn to death. Why? Because if your decision process were to output ‘Left’, then Omega would have predicted that. So there would be no bomb in the box, and you could save yourself the effort of having to clean by taking the Left box[4] - even though that would mean changing the past.
Proponents also sometimes say, by putting the same math into words differently[5], that FDT chooses to take the Left box and burn to death because it imagines that you must be in the reality where Omega made a mistake, and that you can currently affect 999,999,999,999,999,999,999,999 other versions of you and cause them to not have to clean—even though we have no reason to think these other versions exist.
Many people find this kind of recommendation, and either of these framings, extremely implausible—both MacAskill and Bentham’s Bulldog see it as a knock-out blow to FDT, and Wolfgang Schwarz, an academic decision theorist who was a referee of the FDT paper, calls it “insane”. LessWrongers, on the other hand, insist that this way of thinking actually makes sense.[6]
The debate has landed at an impasse—both sides have planted their flags and refuse to budge from their deeply held intuitions.
But I believe that there is a way out.
Many people have made the observation that LessWrong-style decision theory is trying to do something different, in some deep worldview-level way, from academic decision theory—and that this is what leads to the persistent and seemingly intractable disagreement between the two camps.
For example, Paul Christiano:
I think a lot of [academic decision theorists not being interested in LessWrong-style decision theory] is this semantic disagreement / this understanding of “what is the project of decision theory?”…like a difference in, “What are the questions that are interesting, and how should we use language?” Like, “What do concepts like ‘right’ mean?”
I think that there is this (obvious to LessWrongers, because it is deeply entangled with the entire LessWrong philosophy) ontology…that most decision theorists haven’t really considered.
FDT is not playing the same game as CDT or EDT… So it’s odd to have a whole paper comparing them side-by-side as if they are rivals.
So… let’s run with this. Let’s assume that LessWrongers were wrong all along when they thought they were doing decision theory. Let’s come up with a different name for whatever it is that they’ve been doing: call it commitment theory. And let’s reexamine the topic with a fresh eye.
The great commitment
Let’s step back from the extremity of “Bomb*” for a minute—take a simple blackmail scenario:
Omega, a near-perfect predictor, knows a piece of embarrassing information about you. He threatens you that he will release it to the public if you don’t pay him 100$. However, making the threat is slightly costly to him, and he wouldn’t have done it if he hadn’t predicted you would pay.
Everyone agrees that if you can precommit (somehow force your future self) to not pay before you get into this situation, you should—since that would make it so Omega won’t blackmail you and you don’t get into this situation in the first place.
So there’s a class of situations where you would like to previously commit yourself to a choice, to get better outcomes overall. That’s interesting.[7]
So maybe we should do a kind of “generalized precommitment” for those situations, where we commit now to always doing whatever it would have been good to have precommitted to?
Yeah—it seems like, by definition, it would be good to do that. It’s not clear yet what exactly it means for something to be “good to have been precommitted to”, but… we can figure out the details later. I mean, this is tautological: It just has to be good to precommit to… the actions that would have been good to have precommitted to.
In fact, we should do this ASAP, before we end up in such a situation. So let’s do it right now, together, on three: one… two… three!
Well done! Let’s call whatever we just did, this “generalized precommitment”, the great commitment.
What does it mean concretely, that we made this “great commitment”? Well, we’ll get better outcomes if we get put into the blackmail scenario above in the future—so that’s good.
What about other cases? What about, say… Parfit’s Hitchhiker?
You’re in the desert, running out of water and soon to die, when Omega, a near-perfect predictor, finds you. You don’t have any money on you—but he’s not going to save you for free. He’ll only bring you to the city if he predicts you’ll pay him 1,000$ upon reaching it. But once he brings you to the city, Omega can’t force you to pay.
Omega brings you to the city. You are standing in front of the ATM.
Do you pay?
In the past, we obviously wouldn’t have paid, since it would lose us 1,000$ for no reason. But now, as commitment theorists, we think:
Oh man… I wish I could save the 1,000$ by just not paying. But… in advance, I probably would have precommitted to paying, since otherwise Omega wouldn’t save me. So I guess this is covered by the great commitment, and I have to pay.
Darn, I really wish I could just leave and not pay… but it just doesn’t feel “right” somehow. And not even in a moral sense, just that… I can’t imagine doing it? Even though it’s obviously irrational. I guess this is just the kind of guy I am, ever since I made the great commitment.
Actually, I guess that’s why I’m even here and Omega saved me in the first place. Phew, thank God I happened to make the great commitment before I went into the desert.
Okay, that’s interesting. It seems like this might be pretty useful!
Commitment theory
But what does this all mean? What did we just commit ourselves to, exactly?
That’s exactly what commitment theory research consists of—the study of what situations are covered, and how much, by our vague intuitive notion of the “great commitment”.
And since it seems like the great commitment is tautological and like all agents should make it, commitment theory research seems like an important part of the study of ideal decisionmaking—even if it may be distinct from decision theory as such. One might say that decision theory is concerned with what choices are rational, and commitment theory with how, precisely, we should force our future selves to be irrational.[8]
FDT
So how does this relate back to LessWrong-style decision theory?
Well, you might have guessed where this was going already—sorry for tricking you into it. When you made the great commitment earlier, you actually, for all intents and purposes… became an FDT agent[9]. Congratulations!
Yes—commitment theory gives the same practical recommendations as FDT[10], seemingly without requiring any big metaphysical claims about changing the past or alternate realities[11], or any big normative claims about seemingly absurd choices actually being “rational”.
(The only normative claim is that you should make the great commitment ASAP, if you haven’t already—again, it’s tautologically good!)
This also finally gives us a way of reinterpreting some of the complicated technical machinery[12] that people have come up with over the years in their study of LessWrong-style “decision theory”—all this kind of stuff:
Wei Dai (2010) fiddling with the implementation of the concept of “updatelessness” (do what you would have precommitted to).

Macé, Clifton and Kollin (2023) trying to make updatelessness just a little bit updateful, so it isn’t stuck with your past self’s uninformed precommitments (although it probably doesn’t end up working).
In effect, it’s exactly the commitment theory research that I referred to before: trying to pin down and formalize our intuitive notion of “what would have been good to have precommitted to”—trying to figure out what exactly the great commitment covers.
A detailed example
Lastly, let’s reexamine “Bomb*” from our new perspective:
You’re alone at home at the end of the universe. Omega—an amazing predictor of human behavior, but long since dead—gave you one last challenge. You face two boxes, Left and Right. You must take one of them.
In the Left box, Omega put either nothing or a live bomb; if there’s a live bomb, taking this box will set it off, setting you ablaze, and you will burn slowly to death. In the Right box, Omega put a glitter bomb; taking this box will set it off, covering you and your whole living room in glitter, and you’ll have to spend hours cleaning it up—very inconvenient.
Both boxes are transparent, so you can actually see the contents of the Left box—there is a live bomb.
Omega predicted whether you would choose Left or Right when faced with exactly this situation—of seeing a bomb in the Left box. If he predicted that you would choose Right, he put a bomb in Left (like you’re seeing now). If he had predicted that you would choose Left, he wouldn’t have put a bomb in Left, and the box would be empty.
You know that Omega has a failure rate of one in a trillion trillion in this situation.
You are the only person left in the universe. You have a happy life, but you know that you will never meet another agent again, nor face another situation where any of your actions will have been predicted by another agent. What box do you choose?
As a commitment theorist, your thoughts in this situation might now go along these lines:
OH MY GOD THAT’S AN ACTUAL BOMB. Holy fuck. Am I going to die??
Jesus Christ. Oh my god. Why would Omega do this?? Oh my god.
I can just set off the glitter bomb, right? I don’t care about the cleanup, I just don’t want to die!!
But wait—Omega was predicting me, as usual, so this is definitely covered by the great commitment. So I need to act correctly.
Okay, okay, let’s stay calm—what is actually going on here? What would my past self have precommitted to?
(you take a deep breath)
Well, obviously I’d like to take the Right box—if I had committed to that and Omega predicted it, what would have happened? There would be a live bomb in the Left box, I would take the Right box like it predicted… and I’d have to do a few hours of cleaning. Okay, that’s not that bad.
What if I had committed to taking the Left box? Omega would have predicted that, there’d be no bomb in the Left box, and… I could just take the empty Left box and I wouldn’t have to clean.
…
(you start feeling a knot in your stomach)
Wait, wait, wait—there is a one in a trillion trillion chance of Omega being wrong. So for the Right box commitment, there’d actually be a one in a trillion trillion chance of getting an empty Left. And for the Left box commitment, there’d be a one in a trillion trillion chance of actually getting a bomb in the Left box, which I would have committed to take, and burning to death.
So, summing up, from my past self’s perspective: Right box, ~100% chance of having to clean up for a few hours. Left box, one in a trillion trillion chance of burning to death.
That means… the question is if my past self would have valued not having to clean over a one in a trillion trillion chance of burning to death.
(tears well up in your eyes as you start realizing)
I’ve always been a bullet-biting utilitarian type of guy.[13]
One in a trillion trillion is extremely small.
(the tears are streaming down your face now)
I… I don’t want to die.
But there is a bomb in the Left box—that means… Omega predicted I would choose Right in this situation.
What?? But why??? I just said that I wouldn’t do that!!!
…why am I here?
Did Omega make a mistake? Is this a prank? But no, I know that he doesn’t lie.[14]
(screaming at your empty room in disbelief) OMEGA!!! I know you’re long dead but why did you think I would choose Right???? Why????
(you collapse on the ground and cry)
(20 minutes pass)
I don’t want to, I don’t want to, I don’t want to.
You incompetent asshole. Why couldn’t you just be perfect at predicting me, like you usually are? One in a trillion trillion. I really had to get that unlucky.
…
(you stare at the bomb and imagine it exploding as you take the Left box, drenching you in flames)
(you think about how burning to death is said to be among the most painful ways to die)
God, I’m so scared, I’m trembling. It’s so obvious to me now. The incommensurabilists were right all along. This is worse than any amount of cleaning could be good. I was such an idiot. Such a naive, utilitarian moron.
…I don’t actually believe that. I’d be happy with my choice if I had gotten luckier and didn’t have to clean.
(you turn to look at the Right box with the glitter bomb)
Maybe…? I could just…?
(a wordless desperation runs through you)
(you lunge forward to grab the Right box—but something stops you, with your hand only inches away)
…I can’t.
(you drop your hand, sadly)
It’s too deep in my bones. I have to act correctly. I’ve done it for so long. I can’t betray myself like that.
It doesn’t make any sense, but… I could never live with that. I would hate myself.
Sometimes you just get unlucky. That doesn’t mean you get to betray your past self.
…
Yeah. There’s an honor in that. A righteousness.
(you pace around for another hour, reflecting on your life, mourning)
(finally, you go to stand in front of the Left box, and take a final deep breath)
(the tears on your face have long since dried)
(you take the Left box, the bomb explodes, and you burn to death)
Where FDT was doing weird stuff, commitment theory makes the equivalent decisionmaking seem more understandable. Of course, it remains counterintuitive—as Bomb* was designed to be maximally counterintuitive. But hopefully this makes it at least conceivable that a real person might actually take the Left box.
I do believe the great commitment is possible. Things like this are normal: for example, you might tell yourself to get up early tomorrow—and that itself makes you to do it, even if you don’t feel like it in the morning. Human brains are physical systems and, via internal physical manipulations (cognitive activity), can affect their future mental state.[15]
In my mind, it’s a psychological change, some deep realization of a new, non-decision-theoretic notion of “acting correctly”—something like “cooperating with”, being “loyal to” and “coherent with” your past self.
Whether it’s psychologically possible for you to follow such a principle as far as in Bomb*, when the stakes are life-and-death, is an empirical question. But that doesn’t change the fact that the situation is clearly covered by the great commitment, i.e. that you would want to precommit to choosing to burn to death in Bomb*, if you could.
And… you can.
Conclusion
Here’s what I think your takeaways from this post should be:
If you are an adherent of academic decision theory: Consider whether the commitment theory framing makes sense, and whether you should adopt it. It’s a way for you to rescue something useful out of LessWrong-style decision theory[16], while avoiding any of its facially ridiculous claims. Consider whether to make the (tautologically beneficial) great commitment—whether to form a psychological attachment to “acting correctly”, ideally one strong enough that your future self really would choose the Left box in Bomb*.
If you are an adherent of LessWrong-style decision theory: Consider whether to try to convince people of commitment theory instead of FDT—it’s much more contained and much less counterintuitive[17]. It’s also easier to convince people that there’s a new consideration that they just haven’t thought about yet, than that the way they’ve been thinking about something familiar is wrong. Commitment theory is as close to your full vision as people will get within a “normal worldview”—and it’s quite close, since they will agree with you on all actual decisions. You should be pretty happy with that.
- ^
To be clear, it’s not my personally preferred framing, and in my opinion it ultimately fails because no one can actually avoid metaphysics, and LessWrong thought is much closer to correct—but I think it goes quite far, as I’ll argue in this post.
- ^
Those about updatelessness, specifically, not those about counterfactuals (which are already present in academic decision theory).
- ^
I’ve rewritten it to make it clearer, since the original is a bit confusing and underspecified (see this Stuart Armstrong comment), and including that Omega specifically simulates you to predict you is unnecessary. My version preserves the structure and should meet the core desideratum of its proponents (that FDT chooses to burn to death for seemingly no reason) - but maybe call it Bomb* (with a star at the end) for clarity, if you reference it.
It’s kind of like the Transparent Newcomb’s Problem, but optimized for counterintuitiveness—for example, that you are in an “impossible reality” and not in a “possible reality that you just need to make real”.
- ^
at least with extremely high probability. (also, for convenience this paragraph is taken from Will MacAskill and modified for my context—I couldn’t find a better way of explaining it than his, but I also needed to make too many small changes to make it a direct quote.)
- ^
This is a more UDT-style framing.
- ^
e.g. there are a variety of arguments in the comments on Will’s post.
- ^
This is the academic discussion on “binding” and dynamic choice.
- ^
One might also gesture to some intuition that this is central to what “being an agent” means—you only have the ability to make plans if you can lock your future self into a set of actions, if you can expect your future self to “cooperate” with your current self’s commitments. Decision theory and commitment theory seem like two sides of the coin of ideal decisionmaking—how to maximize for your preferences (an adversarial position towards the world) versus how to intentionally ignore and blind yourself to your preferences (to cooperate with the world). (This framing is inspired by Richard Ngo).
- ^
This is the famous “Non-FDT agents self-modify to be FDT agents” argument, put into practice.
- ^
This is one of the core ways that UDT was originally characterized. The clearest statement of that I’ve been able to find is Demski and Garrabrant (2019), who straight-up define it as: “UDT [recommends] that the agent do whatever would have seemed wisest before—whatever your earlier self would have committed to do”.
(and FDT is an umbrella term for “UDT-ish approaches to decision theory”. Although there are some terminological subtleties—there’s a longer comment by Rob Bensinger in that post that’s quite clarifying. But it also comes up several times in the FDT paper as a core property, just ctrl-F “committed”.)
A key point is that humans (and likely almost all agents in the universe when they realize that they should make the great commitment) face basically no Son-of-CDT issue—there is no one that has been predicting their pre-great-commitment self with enough fidelity for it to matter. (It would have to be high enough fidelity to distinguish between making the great commitment and adopting FDT—hard to imagine, since they are so similar).
Of course, you still need the correct counterfactuals—e.g. when deciding what to precommit to in a situation with an agent that is similar but not identical to you. But counterfactuals are the “decision theory dimension” (EDT vs CDT), not the “commitment theory dimension” (updateless vs updateful), so they are a separate issue. (But I think e.g. commitment theory EDT just works).
- ^
It’s buried pretty deep in the post, so I’ll quote it—Christiano says: “I’m not sure if I’m thinking about worlds that don’t exist, or if it’s us who don’t exist and there is some real world somewhere thinking about us.”
- ^
Not all of it—Some of it is better classified as EDT vs CDT stuff (i.e. the “decision theory dimension”), and some as updateless vs updateful stuff (i.e. the “commitment theory dimension”).
- ^
If this makes you disconnect from the scenario because your past self wouldn’t be this utilitarian, just imagine some other less bad consequence. Something like this works no matter what your preferences are—the point is just that FDT can make you choose a locally bad outcome for seemingly no reason.
- ^
Of course, if you face such an unlikely situation in real life, you should completely rethink your assumptions and seriously wonder if you’ve gone crazy. But as it’s a hypothetical decision problem, we’ve fixed the agent’s ex hypothesi beliefs about the situation he’s in—so I’ve written this snippet like you just happen to not think of that possible response.
- ^
Once again, this is the academic discussion on binding.
- ^
I want to briefly examine other attempts to do something similar, and explain why I think they fail. First, Will MacAskill’s “Global CDT”, as it’s closest in the social graph:
Even as it succeeds at replicating a lot of FDT’s recommendations, it leaves less of its conceptual attractiveness intact. Global CDT is framed as simply evaluating with CDT many different “evaluative focal points”: different personalities, dispositions, rules, and so on. But we’d prefer to retain the intuitive sense that the “FDT insight” is one singular thing (“act as you would have precommitted to”), and still avoid the associated metaphysical and normative claims—which we can, with commitment theory.
More importantly, I think it fails on counterfactual mugging with a deterministic coin, since the reality where the coin lands heads is not real and could never have been real, and so doesn’t get taken into account in any updateful decision theory after you learn it lands tails (even if we evaluate bigger things like dispositions). So for situations that we didn’t anticipate, and couldn’t have taken into account when we formed our dispositions or rules or plans, if we aren’t updateless commitment theorists we will fail. (You can understand dispositions and rules as leaky human approximations of theoretical “true” updatelessness, which just evaluates anew the correct precommitment to have taken for every situation).
For attempts from academia (this is going to be uninteresting if you’re not in the weeds on this stuff, so feel free to skip, and I haven’t fully read any of these papers so take it with a grain of salt): I expect the worries above to also apply to Fisher’s and Gauthier’s disposition-based decision theories, Parfit’s “rational irrationality”, McClennen’s “resolute choice” (“it is rational to follow through on plans even when they become locally harmful”) and Bratman (Intentions, Plans, and Practical Reason, 1987). Spohn (”Reversing 30 Years of Discussion”, 2012), Poellinger (“Unboxing the Concepts in Newcomb’s Paradox”, 2013) and Hedden (”Counterfactual Decision Theory”, 2023) are about counterfactuals (the “decision theory dimension”), and not updatelessness (the “commitment theory dimension“), so they’re orthogonal to this discussion.
Meacham’s cohesive decision theory (“do as you would have bound yourself to do”, quite similar) is the closest academic relative to FDT (footnote 34 is incredibly fascinating and prescient as an early discussion of the idea of updatelessness) and also has seemingly the exact same recommendations (modulo a theory of counterfactuals) - but as it also has acts as its evaluative focal point, it incurs the same intuitive drawback of calling it “rational” to choose to burn to death in Bomb*. Also, since it lacks the intuition pumps I’ve built up around it, and tries to fit it into the unnatural frame of decision theory, it ends up looking a bit unmotivated (which is maybe why it didn’t get much traction).
- ^
Especially because, in my opinion, you inevitably run into metaphysical questions (changing the past, alternate realities, platonism about computations) if you really try to justify why e.g. taking the Left box and burning to death in Bomb* is the “correct” decision, in the decision-theoretic sense of maximizing your utility. (Not sure how much of a hot take that is, I will justify it more in a future post). And it’s understandable that people are reluctant to wade into that.
Finally, a brief and dense sidenote (don’t worry if this doesn’t make sense to you): Commitment theory is actually slightly better than naive TDT-style FDT (which is the way many people conceive of FDT), because it performs better on counterfactual mugging with a deterministic coin. This is the sense in which FDT lacks a theory of anthropics, in the “UDT = FDT + a theory of anthropics” way of carving up the space—another way of seeing that metaphysics comes into play.
It seems like you’re asking us to retreat on the fight for the meaning or content of “rational”, in order to “make its concrete recommendations much more palatable to people”, but I feel like winning the meaning/content of “rational” is actually much more important than the latter (in part because I have little idea what concrete recommendations LessWrong-style decision theory even has).
I’m also pretty confused what “rational” really means or implies, but as a starting point, it (alongside values or morality) is a core part of “normativity” (in other words, shouldness, or what’s right), and it really doesn’t make sense to me to concede that “we should force our future selves to be irrational” before we’ve definitively ruled out a much more elegant understanding of “rationality” such that we should simply do what is rational at all times (which can be seen as a core goal of LW-style DT).
Hey Wei, thanks for your work over the years.
Yeah I definitely agree that could be the case, I definitely wouldn’t want LessWrongers to adopt this framing as their actual model (and as I say in footnote 1 I don’t think it ultimately works, and LessWrongers are closer to correct).
But I do feel that as a communication strategy (not in the sense of PR, in the sense of one-on-one conversation), it can often make sense to talk about smaller claims first instead of the deepest cruxes immediately. It can make future conversation about the deepest cruxes easier once we have some appreciation for the other person’s perspective about the smaller claims. E.g. in the fantasy world where everybody internalized what I’m saying here, they would stop disagreeing about Newcomb’s Problem or Bomb* and start discussing metaphysics
and ECLinstead. That would be a positive effect!Could you elaborate on the point in footnote 11 about why you think it doesn’t work? In any concrete problem (Newcomb’s paradox, prisoner’s dilemma, XOR blackmail etc) it seems like all that’s needed is to imagine you made the precommitment at any time before you first had an experience that made you know you’d be facing this problem and need to make a choice (eg before you got the letter in XOR blackmail, but not necessarily before your house was or wasn’t infested with termites since that fact is unknown to you until after getting the letter and making the choice). In what circumstances would you need to think about going further back than that, or some kind of “choice made outside time” perspective? When you worry that agents “might already be in the midst of executing strategies that they might together have precommitted to not doing in advance”, is this not about concrete problems that are sharply bounded in time and have a finite set of well-defined choices like the examples I mentioned, but more like attempts to apply decision theory to ongoing situations that are less well-defined, eg tactics of political negotiation and military aggression in an era of mutually assured nuclear destruction?
No, it’s not about complex vs simple situations, it’s just aboutECLspecifically (at least for humans that’s the only son-of-CDT issue).If I’m understanding correctly, if we are thinking about the ECL problem with a one-shot prisoner’s dilemma, the main new wrinkles relative to other problems are that 1) each agent is unsure that the other will actually make the choice recommended by FDT/the great commitment, 2) observing one’s own choice may provide some new evidence about the probability the other agent will make a given choice (probably modeled as a Bayesian update, which some might think of as an ‘acausal influence’), and finally 3) the original FDT/great commitment recommendation has to take all this into account somehow.
It seems likely to me that one could analyze such a problem in an EDT framework with pre-commitments, where both agents have some probability of backing out of the pre-commitments, and don’t know for sure what those probabilities are so they have to estimate them in making their choice. But in such an analysis I wouldn’t think there’d be any need for consideration of the “metaphysical” issues you talk about in footnote 11 about considering the past long before entering this particular prisoner’s dilemma problem. So is your view that such considerations would be needed to deal with the issues 1), 2), 3) I listed above, or do you see separate issues in an ECL prisoner’s dilemma problem?
Hm, okay I think you’re right actually. Thank you for prodding at this. I mean, it’s not even clear what distinguishes ECL from the prisoner’s dilemma with a twin (modulo counterfactuals). So I think you’re right that updatelessness isn’t needed there, and you only need the correct counterfactuals (the EDT vs CDT dimension), as you said. So there’s no son-of-CDT issue at all for humans (which is actually great news for commitment theory!). I’ll rewrite that footnote.
I think updatelessness can be important for ECL.
If we were to try to go updateless “all the way” and unupdate on a ton of information we’ve already learned, this could have pretty intense implications for ECL. For example, we might no longer have any particular reason to believe that our own values are common across the universe[1] (insofar as our beliefs about this are largely based on our observations), so when deciding what values to optimize for, we might give more weight to other values.
And the case for updatelessness is pretty strong if you endorse EDT, given that EDT with
updatelessnessupdatefulness double counts. I think this is especially true for empirical updatelessness, but also somewhat true for logical updatelessness. (See my comments on that post.)Or that our values are especially correlated with our decision algorithms, which matters a lot for ECL.
Thanks for finally making me read that Paul Christiano post in depth—extremely interesting. I’ll need to digest it for a while.
(until then, you may be interested in my most recent post, where I try to give a metaphysical justification for something like updatelessness)
Okay, so I think you’re right that strong updatelessness like that can change the implications of ECL (although I guess I didn’t technically say otherwise, I only said that updatelessness isn’t needed to get to a version of ECL—and of course, commitment theory isn’t my actual position anyway).
You could see my post that I linked above as an argument in a direction similar to Christiano’s point too—that the possibility of being a freak observer makes you revert back to your ~prior over and over again with every decision (since all your memories could be fake). It’s kind of the very elaborated/general (see especially footnotes 10 and 11) version of this brief digression in one of your comments: “In large worlds where your observations have sufficient randomness that observers of all kinds exists in all worlds, the SSA update step cannot exclude any world. You’re updateless by default.”
However, my current general attitude on a lot of this stuff is that I don’t understand why the decision-theoretic perspective of “how updateless should we be” is a better framing than the metaphysical/ethical perspective of “what should we think of as existing/mattering”. In small worlds with no copies of you (i. e. rejecting one of Christiano’s assumptions), EDT with updating intuitively does fine (or at least doesn’t double count). So the crux here feels less productively described as “how updateless to be” than as our metaphysics/ethics.
For example: “If we go updateless all the way, we might no longer have any particular reason to believe that our own values are common across the universe”. This feels like a confused way to point at the underlying concern (of how to sum up / weigh all the possibilities, which I share) - and confused precisely because it avoids the metaphysical implications. It would be more accurate to say that if we go updateless all the way, we don’t even know what kind of universe we’re in in the first place, if it’s even large and whether ECL and the double counting argument is justified in the first place, etc. (footnote 11 might be relevant here again). If you’re restricting yourself to “the universe as we know it”, you’re not really going “updateless all the way”—so of course this will lead to absurdity (like the conclusion that we can’t get any information about this universe based on our observations in it). (More accurately, it’s selectively updateless—updateless about the existence of our entire civilization, our values, decision processes, etc—but not about the information of our universe being very large that we got from our human community of physicists. That seems nonsensical.)
What do you think? I’m pretty unsure here, definitely still feel confused. Feel free to only respond briefly or to part of it. (also, I left a brief comment on the Christiano post if you’re interested).
(sorry for the nitpick but I think you meant to say EDT with updatefulness double counts?)
you’re right, thanks, edited
Yeah, I also changed the ECL example screenshot. This is much better, thanks a lot.
Oh, and I actually wouldn’t even say that I’m centrally asking you to retreat from the meaning of “rational”. I think if you framed it like “there are two facets of rationality, the decision theoretic one and the commitment theoretic one, and you just hadn’t realized that there was a non-decision-theoretic one until now”, that would be very similar.
Do you think it makes sense to think about rationality this way, that it has two facets, aside from “palatability”? I guess my key consideration is wanting to optimize for conceptual clarity/elegance (or get an argument for why that’s not possible or desirable), and not wanting to compromise on that for the sake of convincing people to take LW-style DT more seriously, even in a one on one setting, unless it’s pretty easy to walk back the compromise later and ultimately get everyone on the same page about the real concept. (Otherwise, we get more uptake in the short run, but more conceptual fragmentation/confusion in the long run, which I don’t think is a good tradeoff.)
Have you tried having conversations with people this way, and how did it work out?
Well, we are thinking in terms of two facets already if we are thinking e.g. in terms of the (EDT vs CDT)x(updateful vs updateless) 2x2! (which is just decision theory x commitment theory in the framing in the post). So I don’t think that’s that big of a deal.
If you permit me some speculation, is it possible you’re not actually worried about the meaning of “rationality”? (since it’s unclear to me why it having two facets would be bad). But instead about the minimization of metaphysical claims like “you are your algorithm”, the possibility of alternate realities, etc, that I’m doing in the post? Maybe it’s just because I’m working on another post about how people seem to sweep their metaphysical cruxes under the rug in decision theory debates and I’m just seeing that everywhere, but do you think that’s possible?
(and I agree with these metaphysical claims, and I think it’s valid to be worried about them being minimized, since they might be important, and there’s a value of “conceptual correctness/clarity” there too—but it’s different from being worried about the concept of rationality being diluted)
(Haven’t had any conversations sadly, am not in a location irl where there are many people who want to talk about decision theory).
The distinction between changing and controlling/determining/deciding is rather important to how this works. There clearly isn’t any changing of the past, “change” doesn’t seem like an interchangeable term for describing what’s going on. Carlsmith’s post you link to has things to say on this exact point:
I don’t disagree with that intuition, but I don’t think you can get away from the fact that you’re making a metaphysical claim. You’re saying that e.g. what the person in Bomb* is seeing is (almost certainly) not real, in some sense. That’s getting metaphysical. I don’t think there’s a way around that.
I’m objecting to the use of the word “change”, specifically.
Ah, gotcha—but on a pre-theoretical, intuitive level, it just is “changing” the past—since you trust your seeming direct perception of reality.
The past that was perceived also isn’t being changed. It stays as it was perceived, in that reality. There is no claim that it magically changes. (What happens is that your decision can determine that reality’s weight/probability to have always been negligible compared to the weight of the other relevant possibilities.)
OK, so this is a more UDT-style framing (which I’m sympathetic to). Then the “alternate realities” part, not the “changing the past” part, was more the one targeted at you. But there are definitely people who think about FDT as “changing the past”, or at least talk about it in that way, so that’s who that clause was targeted at.
I have an alternative ending:
I’m in favor of this reframe, and it gels well with its real world applications. In fact, Zvi wrote a post a long long time ago for which commitment theory would be a more convincing name than decision theory:
https://www.lesswrong.com/posts/scwoBEju75C45W5n3/how-i-lost-100-pounds-using-tdt
Centralizing the “negotiates well with similar agents / future selves” framing is actually pretty intuitive in certain everyday situations!
Thanks a lot Rachel!
That Zvi post is a classic, thanks for reminding me of it.
Although I would actually make a distinction here—“negotiates well with similar agents” is just about your counterfactuals, so commitment theory / updatelessness is not necessary for it. e.g. updateful CDT with logical counterfactuals and updateful EDT both are able to handle that. So you can just use normal academic EDT.
Negotiating with your future/past self, that’s where it gets thorny. That’s what you need commitment theory / updatelessness for.
I would have liked a warning before this phrase, because as I was reading it, one part of my brain automatically started thinking about trying to precommit to the Great Commitment, while another was (literally) screaming very loudly that it would be extremely stupid to do so without thinking it through. I was just left confused for a bit.
Of course, this kind of precommitment, if it is possible at all, would not arise simply from reading the phrase and trying to parse it. But the experience of going through that without being warned beforehand was still a bit disturbing.
(If you want to immediately skip to why I believe past attempts to incorporate LessWrong-style intuitions into standard decision theory are insufficient, go to footnote 16)
Doesn’t the traditional lesswrong Bayesian Utilitarianism save you from Omega’s bomb? Whatever decision theory you believe in, you must have to it in order to justify left boxing. Since there’s no rigorous basis to assign probabilities to decision theories, your probability can’t be that high and you end up right boxing.
I believe they way it works is that FDT tells you to make the bad decision (here: suicide) if faced with the actual situation, but the argument is that you won’t get into the situation at all or less often because you’re playing FDT.
By the time you find yourself facing the problem in which FDT recommends killing yourself something’s already gone wrong, because the situation ought to be prevented by playing FDT in the first place.
But of course the question always is: should you actually find yourself in the decision problem, despite being an FDT-agent that ought to never get here, should you still follow FDT?
I’d argue the answer is “lol no” but I believe the argument in favor is to assume you’re just “hypothetical you”, evaluating the situation without having to bear the cost. E.g. you’re just a simulated version omega spon up to check how you’d react in this decision, not “actual you” who’s faced with the decision. After all “actual you” is not supposed to even get into the situation. If you truly believe that, then chances are you’re not you, because that’d mean omega is wrong, which is extremely unlikely.
Yeah, I think this is a good argument to give people an intuition, but it doesn’t answer the question of “OK, but if you are in the situation, what do you do and why?”. FDT supporters (of which I am one, to be clear) do need to answer that.
I though the answer was “assume you’re simulated and follow FDT to save real-you the trouble of ever getting into this situation in real life”. Out of the all arguments to take the bomb this is the only one I’ve ever heard which I can at least understand where it’s coming from.
If that is not the FDT response then, I guess I don’t know why you’d ever really blow yourself up. I did read the whole post, including rereading the example against just now.
But what I got from it was mostly the insight that FDT kind of answers a different question than CDT, in that it’s goal is to shape what situations you end up in, not necessarily how to get the best result out of a given situation.
Thanks for the post—it gave me plenty of food for thought!
On the object level: IMHO, asking the question “OK, but if you are in the situation, what do you do?” follows from an implicit rejection of the scenario’s embeddedness — and without embeddedness, FDT doesn’t make sense. I’ve written a 3-min post elaborating on this.
did i just precommit to always saying $9 in a $10 ultimatum game?
i guess it would be good to have done so...