So the whole direction of justifying probabilities by decisions is probably a dead end.
CDT+SSSA+SIA gives KKT conditions, covers all UDT policies in absent-minded games, is ex ante optimal in additive games.… A bit more discussion in other post. The part where number of copies is affected does not prevent CDT+SSSA+SIA from being well-defined (in the sense of having self-ratifying probability / policy combinations), or from giving ex ante utility stationarity conditions. It just complicates the scenario, which is why I didn’t focus on it in this post.
Alternatively, if you view probabilities as “reality fluid” that determines what to expect in the next moment, that’s an important question too, but it’s not clear why Dutch books help solve it.
This seems more in metaphysics territory. I take the physicalist baseline to be that one should assign probabilities to centered worlds and that’s all. Maybe there is other stuff to assign probabilities to (like qualia, idk). I think before assigning numbers a thing to do is to figure out what is the metaphysical status (if any) of the entity being assigned probabilities. Since it does not seem clear to me that there is a there there. (Whereas with things like the Sleeping Beauty problem or Doomsday Argument it seems there is plenty of disagreement to be had about rational probabilities on centered worlds, even without getting into any metaphysics on top of centered worlds)
Thank you for the links! I’m not sure they fully answer my question though. It’s more about what we’re trying to do.
Maybe we want to start with probabilities and derive the optimal policy, continuing the project of VNM utility maximization. To me that project basically died with Absent-Minded-Driver, because it shows we can’t start with probabilities. Having “self-ratifying probability/policy combinations” doesn’t have quite the same pull. And in any case the necessity of self-coordination (UDT1.1) forces us to choose whole policies, not individual actions. Or do you hope that if we push probability-based approaches far enough, they can close the gap?
Or maybe we first choose which policy to follow, then compute the probabilities of finding ourselves at one node or another. That’s fine, and I do think SIA is the best answer. And maybe even we can show that these probabilities have nice decision-making properties, but to me this seems like an “echo” of us choosing the best policy to begin with, no?
Or maybe I’m being silly again and misunderstanding the whole thing?
Primarily 2, with the extra bit where we can be assured there’s a CDT+SIA fixed point. For the purpose of anthropic probabilities, that’s a desirable condition. We don’t need probabilities to force optimality but we would rather they be compatible with optimality, and preferably that they help with the computation.
The other part is that UDT is not really computationally tractable and so figuring out how to optimize it is useful, which is more (1) territory. For this the computational complexity paper is relevant because it shows CDT+SIA (which they call CDT+GT) gives KKT conditions. So we can restrict to CDT+GT points when searching for ex ante optimal policies.
One idea also is to relax to “list CDT+GT policies but without the inequalities and allowing complex numbers” which is now algebraic geometry territory (solving complex polynomials), and then filter for inequalities, real numbers, and global optimality afterwards.
From reading through the Sleeping Beauty literature, I learned about the Wichardt 2008 counterexample for multi-player ex ante optimality, which is relevant to UDT (and says why we shouldn’t expect nice things for UDT in general).
To give a brief summary, suppose Alice has 2 copies, who have the same source code and who can randomize independently. There are 2 coffee shops that the copies can decide to go to without communicating. They would really like (+5) to meet at the same coffee shop. Also, Bob hates Alice, his utility function is hers negated. Alice gets −1 utility (and Bob +1) if Bob goes to the coffee shop where both Alice-copies go. (Basically Alice is playing a “coordination” variant of absent-minded driver with herself, while also playing matching pennies with Bob)
The problem is that they can’t both select ex ante optimal policies in the Nash equilibrium sense. Alice needs to either 100% go to coffee shop A or B to be ex ante optimal; suppose it’s A. But then Bob needs to 100% go to coffee shop A to be ex ante optimal. Now Alice isn’t ex ante optimal anymore because it would have been more optimal for both her copies to go to coffee shop B.
If we treat it as a 3 player game then we have a Nash equilibrium (both Alices go to A, Bob goes to A). This is CDT+SIA optimal on everyone’s part (the SIA part doesn’t matter for these purposes). Since neither Alice-instance can unilaterally switch and expect things to (causally) work out well.
Given this impossibility result it’s not clear what UDT-like decision theories should do in these sorts of situations. Since it means not everyone can be ex ante optimal and maybe we need weaker conditions too.
I see, thanks! The Wichardt example is really interesting. I already kinda knew that UDT in multiplayer games doesn’t work too well (even in a simple 2-player asymmetric game with nonzero sum, equilibrium selection becomes a problem), but the nonexistence of Nash equilibria is even more fun. Would you say that CDT+SIA optimality works better than UDT for multiplayer games, and in how much generality?
I do think “multi player equilibria always exist” is a pretty desirable property for an equilibrium concept, although the lack of ex ante optimality is an issue. (Some of my previous writing on this: “Buridan’s ass in coordination games”, “Reducing collective rationality to individual optimization in common-payoff games using MCMC”, both making use of shared randomness, a workaround for Wichardt 2008 style counterexamples; as an offhand idea, it might be interesting to examine how logical inductor EDT handles such cases.)
Absent minded games can provide “easier” models of Newcomblike games; quoting a previous post,
Memoryless Cartesian environments can be used to define many familier decision problems (for example, the absent-minded driver problem, Newcomb’s problem with opaque or transparent boxes (assuming Omega runs a copy of the agent to make its prediction), counterfactual mugging (also assuming Omega simulates the agent)). Translating a decision problem to a memoryless Cartesian environment obviously requires making some Cartesian assumptions/decisions, though; in the case of Newcomb’s problem, we have to isolate Omega’s simulation of the agent as a copy of the agent.
Which is part of why I’m focusing on them (as they seem less general / easier to analyze).
Yeah, agree on the point about absent-minded games, I think I realized it in 2014. But I didn’t make the jump to multiplayer absent-minded games. It’s cool that you explained it to me now, I’ll spend some time thinking it through.
Maybe the more general problem is I tend to abandon whole directions of thinking pretty easily if I see a “deep enough” problem with them. Abandoning probabilities because of absent-minded driver (as I mentioned in the toplevel comment); abandoning logical induction because of Diffractor’s result that the probability distribution LI converges on doesn’t itself satisfy the LI criterion and can be exploited by traders; abandoning most ideas on equilibrium selection because (edit) there’s just too many with no clear winner. Or maybe it’s a good heuristic and just sometimes works badly, idk.
CDT+SSSA+SIA gives KKT conditions, covers all UDT policies in absent-minded games, is ex ante optimal in additive games.… A bit more discussion in other post. The part where number of copies is affected does not prevent CDT+SSSA+SIA from being well-defined (in the sense of having self-ratifying probability / policy combinations), or from giving ex ante utility stationarity conditions. It just complicates the scenario, which is why I didn’t focus on it in this post.
This seems more in metaphysics territory. I take the physicalist baseline to be that one should assign probabilities to centered worlds and that’s all. Maybe there is other stuff to assign probabilities to (like qualia, idk). I think before assigning numbers a thing to do is to figure out what is the metaphysical status (if any) of the entity being assigned probabilities. Since it does not seem clear to me that there is a there there. (Whereas with things like the Sleeping Beauty problem or Doomsday Argument it seems there is plenty of disagreement to be had about rational probabilities on centered worlds, even without getting into any metaphysics on top of centered worlds)
Thank you for the links! I’m not sure they fully answer my question though. It’s more about what we’re trying to do.
Maybe we want to start with probabilities and derive the optimal policy, continuing the project of VNM utility maximization. To me that project basically died with Absent-Minded-Driver, because it shows we can’t start with probabilities. Having “self-ratifying probability/policy combinations” doesn’t have quite the same pull. And in any case the necessity of self-coordination (UDT1.1) forces us to choose whole policies, not individual actions. Or do you hope that if we push probability-based approaches far enough, they can close the gap?
Or maybe we first choose which policy to follow, then compute the probabilities of finding ourselves at one node or another. That’s fine, and I do think SIA is the best answer. And maybe even we can show that these probabilities have nice decision-making properties, but to me this seems like an “echo” of us choosing the best policy to begin with, no?
Or maybe I’m being silly again and misunderstanding the whole thing?
Primarily 2, with the extra bit where we can be assured there’s a CDT+SIA fixed point. For the purpose of anthropic probabilities, that’s a desirable condition. We don’t need probabilities to force optimality but we would rather they be compatible with optimality, and preferably that they help with the computation.
The other part is that UDT is not really computationally tractable and so figuring out how to optimize it is useful, which is more (1) territory. For this the computational complexity paper is relevant because it shows CDT+SIA (which they call CDT+GT) gives KKT conditions. So we can restrict to CDT+GT points when searching for ex ante optimal policies.
One idea also is to relax to “list CDT+GT policies but without the inequalities and allowing complex numbers” which is now algebraic geometry territory (solving complex polynomials), and then filter for inequalities, real numbers, and global optimality afterwards.
From reading through the Sleeping Beauty literature, I learned about the Wichardt 2008 counterexample for multi-player ex ante optimality, which is relevant to UDT (and says why we shouldn’t expect nice things for UDT in general).
To give a brief summary, suppose Alice has 2 copies, who have the same source code and who can randomize independently. There are 2 coffee shops that the copies can decide to go to without communicating. They would really like (+5) to meet at the same coffee shop. Also, Bob hates Alice, his utility function is hers negated. Alice gets −1 utility (and Bob +1) if Bob goes to the coffee shop where both Alice-copies go. (Basically Alice is playing a “coordination” variant of absent-minded driver with herself, while also playing matching pennies with Bob)
The problem is that they can’t both select ex ante optimal policies in the Nash equilibrium sense. Alice needs to either 100% go to coffee shop A or B to be ex ante optimal; suppose it’s A. But then Bob needs to 100% go to coffee shop A to be ex ante optimal. Now Alice isn’t ex ante optimal anymore because it would have been more optimal for both her copies to go to coffee shop B.
If we treat it as a 3 player game then we have a Nash equilibrium (both Alices go to A, Bob goes to A). This is CDT+SIA optimal on everyone’s part (the SIA part doesn’t matter for these purposes). Since neither Alice-instance can unilaterally switch and expect things to (causally) work out well.
Given this impossibility result it’s not clear what UDT-like decision theories should do in these sorts of situations. Since it means not everyone can be ex ante optimal and maybe we need weaker conditions too.
I see, thanks! The Wichardt example is really interesting. I already kinda knew that UDT in multiplayer games doesn’t work too well (even in a simple 2-player asymmetric game with nonzero sum, equilibrium selection becomes a problem), but the nonexistence of Nash equilibria is even more fun. Would you say that CDT+SIA optimality works better than UDT for multiplayer games, and in how much generality?
While it doesn’t globally “work better”, yes CDT+SIA is guaranteed to have solutions in multi player absent minded games, and UDT isn’t.
There are also EDT+FNC (aka EDT+GDH) equilibrium concepts, which do not always exist in multiplayer games (see “Imperfect-Recall Games: Equilibrium Concepts and Their Complexity”, Lemma 21).
I do think “multi player equilibria always exist” is a pretty desirable property for an equilibrium concept, although the lack of ex ante optimality is an issue. (Some of my previous writing on this: “Buridan’s ass in coordination games”, “Reducing collective rationality to individual optimization in common-payoff games using MCMC”, both making use of shared randomness, a workaround for Wichardt 2008 style counterexamples; as an offhand idea, it might be interesting to examine how logical inductor EDT handles such cases.)
Absent minded games can provide “easier” models of Newcomblike games; quoting a previous post,
Which is part of why I’m focusing on them (as they seem less general / easier to analyze).
Yeah, agree on the point about absent-minded games, I think I realized it in 2014. But I didn’t make the jump to multiplayer absent-minded games. It’s cool that you explained it to me now, I’ll spend some time thinking it through.
Maybe the more general problem is I tend to abandon whole directions of thinking pretty easily if I see a “deep enough” problem with them. Abandoning probabilities because of absent-minded driver (as I mentioned in the toplevel comment); abandoning logical induction because of Diffractor’s result that the probability distribution LI converges on doesn’t itself satisfy the LI criterion and can be exploited by traders; abandoning most ideas on equilibrium selection because (edit) there’s just too many with no clear winner. Or maybe it’s a good heuristic and just sometimes works badly, idk.