Makes sense that probabilities on only outcomes or biases aren’t rich enough to imagine Omega messing with you, but is there some slightly richer probabilistic model that works fine?
E.g. if Omega is predicting you before you even choose and rigging the coin, maybe your hypotheses need to be UDT-style universes that take the “you” program as an input. Learning that Omega is messing with you could be done by updating to place a higher probability on some universes rather than others.
On the one hand, hypotheses about the entire universe are much more extravagant than hypotheses about a single binary variable. But I don’t think it’s crazy that in order to imagine Omega messing with you, you can’t think about the coin in perfect isolation.
That’s a reasonable idea, but if you work through it, you will nonetheless find that if your belief state is represented by a single probability distribution, then when you compute the expected value of Heads and Tails, one (or both) of them will exceed 0.45.
Of course, you could say that your probability distribution is about Omega’s policy, and that it puts all (or most of) its mass on “Heads iff I bet Tails”, and then you can say that your payoff, instead of being an expected value, is some richer pairing of your policy with Omega’s policy. This is roughly the orthodox LessWrong way of handling Newcomblike problems, which is also a departure from orthodox Bayesian decision theory. But then there is still no consistent way to formalize your epistemic state about the coin’s outcome as a probability distribution about the coin’s outcome, even though there clearly is a rational epistemic state to have about the coin’s outcome.
The domain of a probability distribution can be any set whatsoever, and the distribution itself is a measure over a sigma-algebra of subsets of that domain. It’s common for it to be a nice set like the real numbers or some subset thereof, or maybe tuples of real numbers, but nothing requires that.
In particular, it is perfectly valid for one element of that domain to be a world-model such as “the coin is biased to always land opposite of your bet (and if you didn’t make a bet it will always land tails)”, as well as many others. The domain is even allowed to include elements such as “none of the above”. The distribution can then express your credence that the real world behaves according to one of those world-models (or none of the above).
The usual axioms of probability then constrain your rational credences in various outcomes, such as credence in winning certain bets.
Here is a very simple probability distribution over world models that makes it rational to roll the d20:
P(“the coin is biased to always land opposite of your bet (and if you didn’t make a bet it will always land tails)”) = 1. There is no need for any other elements in the domain, and this is allowed by the axioms of a probability space and definition of distribution.
It is clear that this satisfies the condition of coin bias > 0.05, since the coin’s outcome is deterministic. Consequently P(winning | bet Heads) = P(winning | bet Tails) = 0, but P(winning | roll d20) = 0.45 (by definition of fair d20 and application of probability rules).
I see what you mean, but then the question becomes: what is the formal foundation for what counts as a “world model”, beyond “it’s a string in the English language that I can use to make predictions”? That is exactly the question my thinking about formal epistemology is working toward answering.
(The orthodox Bayesian answer is that a “world model” is nothing other than a single joint probability distribution about all conceivable variables at once, and that any outer “hyperprior” such as the one you named can be “marginalized away”. So no matter how many variables you include, this forces you to, in particular, act in accordance with a joint probability distribution about states and outcomes. Therefore, what I have shown is that the orthodox Bayesian view of “world models” forces you into an irrational choice in some Newcomblike problems.)
But then there is still no consistent way to formalize your epistemic state about the coin’s outcome as a probability distribution about the coin’s outcome, even though there clearly is a rational epistemic state to have about the coin’s outcome.
If there’s a natural way to condense your information about the coin into a single probability, you can do it just as well starting from a distribution over UDT-style universes. Like if you think it should be 50⁄50 because that’s your best guess if you forget the information about your precise action, you can give yourself a low-information distribution over policies and then marginalize over it to see what happens to the coin.
Makes sense that probabilities on only outcomes or biases aren’t rich enough to imagine Omega messing with you, but is there some slightly richer probabilistic model that works fine?
E.g. if Omega is predicting you before you even choose and rigging the coin, maybe your hypotheses need to be UDT-style universes that take the “you” program as an input. Learning that Omega is messing with you could be done by updating to place a higher probability on some universes rather than others.
On the one hand, hypotheses about the entire universe are much more extravagant than hypotheses about a single binary variable. But I don’t think it’s crazy that in order to imagine Omega messing with you, you can’t think about the coin in perfect isolation.
That’s a reasonable idea, but if you work through it, you will nonetheless find that if your belief state is represented by a single probability distribution, then when you compute the expected value of Heads and Tails, one (or both) of them will exceed 0.45.
Of course, you could say that your probability distribution is about Omega’s policy, and that it puts all (or most of) its mass on “Heads iff I bet Tails”, and then you can say that your payoff, instead of being an expected value, is some richer pairing of your policy with Omega’s policy. This is roughly the orthodox LessWrong way of handling Newcomblike problems, which is also a departure from orthodox Bayesian decision theory. But then there is still no consistent way to formalize your epistemic state about the coin’s outcome as a probability distribution about the coin’s outcome, even though there clearly is a rational epistemic state to have about the coin’s outcome.
The domain of a probability distribution can be any set whatsoever, and the distribution itself is a measure over a sigma-algebra of subsets of that domain. It’s common for it to be a nice set like the real numbers or some subset thereof, or maybe tuples of real numbers, but nothing requires that.
In particular, it is perfectly valid for one element of that domain to be a world-model such as “the coin is biased to always land opposite of your bet (and if you didn’t make a bet it will always land tails)”, as well as many others. The domain is even allowed to include elements such as “none of the above”. The distribution can then express your credence that the real world behaves according to one of those world-models (or none of the above).
The usual axioms of probability then constrain your rational credences in various outcomes, such as credence in winning certain bets.
Here is a very simple probability distribution over world models that makes it rational to roll the d20:
P(“the coin is biased to always land opposite of your bet (and if you didn’t make a bet it will always land tails)”) = 1. There is no need for any other elements in the domain, and this is allowed by the axioms of a probability space and definition of distribution.
It is clear that this satisfies the condition of coin bias > 0.05, since the coin’s outcome is deterministic. Consequently P(winning | bet Heads) = P(winning | bet Tails) = 0, but P(winning | roll d20) = 0.45 (by definition of fair d20 and application of probability rules).
I see what you mean, but then the question becomes: what is the formal foundation for what counts as a “world model”, beyond “it’s a string in the English language that I can use to make predictions”? That is exactly the question my thinking about formal epistemology is working toward answering.
(The orthodox Bayesian answer is that a “world model” is nothing other than a single joint probability distribution about all conceivable variables at once, and that any outer “hyperprior” such as the one you named can be “marginalized away”. So no matter how many variables you include, this forces you to, in particular, act in accordance with a joint probability distribution about states and outcomes. Therefore, what I have shown is that the orthodox Bayesian view of “world models” forces you into an irrational choice in some Newcomblike problems.)
If there’s a natural way to condense your information about the coin into a single probability, you can do it just as well starting from a distribution over UDT-style universes. Like if you think it should be 50⁄50 because that’s your best guess if you forget the information about your precise action, you can give yourself a low-information distribution over policies and then marginalize over it to see what happens to the coin.