Yeah I definitely agree that could be the case, I definitely wouldn’t want LessWrongers to adopt this framing as their actual model (and as I say in footnote 1 I don’t think it ultimately works, and LessWrongers are closer to correct).
But I do feel that as a communication strategy (not in the sense of PR, in the sense of one-on-one conversation), it can often make sense to talk about smaller claims first instead of the deepest cruxes immediately. It can make future conversation about the deepest cruxes easier once we have some appreciation for the other person’s perspective about the smaller claims. E.g. in the fantasy world where everybody internalized what I’m saying here, they would stop disagreeing about Newcomb’s Problem or Bomb* and start discussing metaphysics and ECL instead. That would be a positive effect!
Could you elaborate on the point in footnote 11 about why you think it doesn’t work? In any concrete problem (Newcomb’s paradox, prisoner’s dilemma, XOR blackmail etc) it seems like all that’s needed is to imagine you made the precommitment at any time before you first had an experience that made you know you’d be facing this problem and need to make a choice (eg before you got the letter in XOR blackmail, but not necessarily before your house was or wasn’t infested with termites since that fact is unknown to you until after getting the letter and making the choice). In what circumstances would you need to think about going further back than that, or some kind of “choice made outside time” perspective? When you worry that agents “might already be in the midst of executing strategies that they might together have precommitted to not doing in advance”, is this not about concrete problems that are sharply bounded in time and have a finite set of well-defined choices like the examples I mentioned, but more like attempts to apply decision theory to ongoing situations that are less well-defined, eg tactics of political negotiation and military aggression in an era of mutually assured nuclear destruction?
If I’m understanding correctly, if we are thinking about the ECL problem with a one-shot prisoner’s dilemma, the main new wrinkles relative to other problems are that 1) each agent is unsure that the other will actually make the choice recommended by FDT/the great commitment, 2) observing one’s own choice may provide some new evidence about the probability the other agent will make a given choice (probably modeled as a Bayesian update, which some might think of as an ‘acausal influence’), and finally 3) the original FDT/great commitment recommendation has to take all this into account somehow.
It seems likely to me that one could analyze such a problem in an EDT framework with pre-commitments, where both agents have some probability of backing out of the pre-commitments, and don’t know for sure what those probabilities are so they have to estimate them in making their choice. But in such an analysis I wouldn’t think there’d be any need for consideration of the “metaphysical” issues you talk about in footnote 11 about considering the past long before entering this particular prisoner’s dilemma problem. So is your view that such considerations would be needed to deal with the issues 1), 2), 3) I listed above, or do you see separate issues in an ECL prisoner’s dilemma problem?
Hm, okay I think you’re right actually. Thank you for prodding at this. I mean, it’s not even clear what distinguishes ECL from the prisoner’s dilemma with a twin (modulo counterfactuals). So I think you’re right that updatelessness isn’t needed there, and you only need the correct counterfactuals (the EDT vs CDT dimension), as you said. So there’s no son-of-CDT issue at all for humans (which is actually great news for commitment theory!). I’ll rewrite that footnote.
Oh, and I actually wouldn’t even say that I’m centrally asking you to retreat from the meaning of “rational”. I think if you framed it like “there are two facets of rationality, the decision theoretic one and the commitment theoretic one, and you just hadn’t realized that there was a non-decision-theoretic one until now”, that would be very similar.
Do you think it makes sense to think about rationality this way, that it has two facets, aside from “palatability”? I guess my key consideration is wanting to optimize for conceptual clarity/elegance (or get an argument for why that’s not possible or desirable), and not wanting to compromise on that for the sake of convincing people to take LW-style DT more seriously, even in a one on one setting, unless it’s pretty easy to walk back the compromise later and ultimately get everyone on the same page about the real concept. (Otherwise, we get more uptake in the short run, but more conceptual fragmentation/confusion in the long run, which I don’t think is a good tradeoff.)
Have you tried having conversations with people this way, and how did it work out?
Well, we are thinking in terms of two facets already if we are thinking e.g. in terms of the (EDT vs CDT)x(updateful vs updateless) 2x2! (which is just decision theory x commitment theory in the framing in the post). So I don’t think that’s that big of a deal.
If you permit me some speculation, is it possible you’re not actually worried about the meaning of “rationality”? (since it’s unclear to me why it having two facets would be bad). But instead about the minimization of metaphysical claims like “you are your algorithm”, the possibility of alternate realities, etc, that I’m doing in the post? Maybe it’s just because I’m working on another post about how people seem to sweep their metaphysical cruxes under the rug in decision theory debates and I’m just seeing that everywhere, but do you think that’s possible?
(and I agree with these metaphysical claims, and I think it’s valid to be worried about them being minimized, since they might be important, and there’s a value of “conceptual correctness/clarity” there too—but it’s different from being worried about the concept of rationality being diluted)
(Haven’t had any conversations sadly, am not in a location irl where there are many people who want to talk about decision theory).
Hey Wei, thanks for your work over the years.
Yeah I definitely agree that could be the case, I definitely wouldn’t want LessWrongers to adopt this framing as their actual model (and as I say in footnote 1 I don’t think it ultimately works, and LessWrongers are closer to correct).
But I do feel that as a communication strategy (not in the sense of PR, in the sense of one-on-one conversation), it can often make sense to talk about smaller claims first instead of the deepest cruxes immediately. It can make future conversation about the deepest cruxes easier once we have some appreciation for the other person’s perspective about the smaller claims. E.g. in the fantasy world where everybody internalized what I’m saying here, they would stop disagreeing about Newcomb’s Problem or Bomb* and start discussing metaphysics
and ECLinstead. That would be a positive effect!Could you elaborate on the point in footnote 11 about why you think it doesn’t work? In any concrete problem (Newcomb’s paradox, prisoner’s dilemma, XOR blackmail etc) it seems like all that’s needed is to imagine you made the precommitment at any time before you first had an experience that made you know you’d be facing this problem and need to make a choice (eg before you got the letter in XOR blackmail, but not necessarily before your house was or wasn’t infested with termites since that fact is unknown to you until after getting the letter and making the choice). In what circumstances would you need to think about going further back than that, or some kind of “choice made outside time” perspective? When you worry that agents “might already be in the midst of executing strategies that they might together have precommitted to not doing in advance”, is this not about concrete problems that are sharply bounded in time and have a finite set of well-defined choices like the examples I mentioned, but more like attempts to apply decision theory to ongoing situations that are less well-defined, eg tactics of political negotiation and military aggression in an era of mutually assured nuclear destruction?
No, it’s not about complex vs simple situations, it’s just aboutECLspecifically (at least for humans that’s the only son-of-CDT issue).If I’m understanding correctly, if we are thinking about the ECL problem with a one-shot prisoner’s dilemma, the main new wrinkles relative to other problems are that 1) each agent is unsure that the other will actually make the choice recommended by FDT/the great commitment, 2) observing one’s own choice may provide some new evidence about the probability the other agent will make a given choice (probably modeled as a Bayesian update, which some might think of as an ‘acausal influence’), and finally 3) the original FDT/great commitment recommendation has to take all this into account somehow.
It seems likely to me that one could analyze such a problem in an EDT framework with pre-commitments, where both agents have some probability of backing out of the pre-commitments, and don’t know for sure what those probabilities are so they have to estimate them in making their choice. But in such an analysis I wouldn’t think there’d be any need for consideration of the “metaphysical” issues you talk about in footnote 11 about considering the past long before entering this particular prisoner’s dilemma problem. So is your view that such considerations would be needed to deal with the issues 1), 2), 3) I listed above, or do you see separate issues in an ECL prisoner’s dilemma problem?
Hm, okay I think you’re right actually. Thank you for prodding at this. I mean, it’s not even clear what distinguishes ECL from the prisoner’s dilemma with a twin (modulo counterfactuals). So I think you’re right that updatelessness isn’t needed there, and you only need the correct counterfactuals (the EDT vs CDT dimension), as you said. So there’s no son-of-CDT issue at all for humans (which is actually great news for commitment theory!). I’ll rewrite that footnote.
Yeah, I also changed the ECL example screenshot. This is much better, thanks a lot.
Oh, and I actually wouldn’t even say that I’m centrally asking you to retreat from the meaning of “rational”. I think if you framed it like “there are two facets of rationality, the decision theoretic one and the commitment theoretic one, and you just hadn’t realized that there was a non-decision-theoretic one until now”, that would be very similar.
Do you think it makes sense to think about rationality this way, that it has two facets, aside from “palatability”? I guess my key consideration is wanting to optimize for conceptual clarity/elegance (or get an argument for why that’s not possible or desirable), and not wanting to compromise on that for the sake of convincing people to take LW-style DT more seriously, even in a one on one setting, unless it’s pretty easy to walk back the compromise later and ultimately get everyone on the same page about the real concept. (Otherwise, we get more uptake in the short run, but more conceptual fragmentation/confusion in the long run, which I don’t think is a good tradeoff.)
Have you tried having conversations with people this way, and how did it work out?
Well, we are thinking in terms of two facets already if we are thinking e.g. in terms of the (EDT vs CDT)x(updateful vs updateless) 2x2! (which is just decision theory x commitment theory in the framing in the post). So I don’t think that’s that big of a deal.
If you permit me some speculation, is it possible you’re not actually worried about the meaning of “rationality”? (since it’s unclear to me why it having two facets would be bad). But instead about the minimization of metaphysical claims like “you are your algorithm”, the possibility of alternate realities, etc, that I’m doing in the post? Maybe it’s just because I’m working on another post about how people seem to sweep their metaphysical cruxes under the rug in decision theory debates and I’m just seeing that everywhere, but do you think that’s possible?
(and I agree with these metaphysical claims, and I think it’s valid to be worried about them being minimized, since they might be important, and there’s a value of “conceptual correctness/clarity” there too—but it’s different from being worried about the concept of rationality being diluted)
(Haven’t had any conversations sadly, am not in a location irl where there are many people who want to talk about decision theory).