they simply evaluate the goodness of actions in a “weird” way—e.g. what makes actions good is that they’re part of the winningest strategy.
So if I’m an AI that was just created and is instantly blackmailed, in what sense *exactly* is not paying the “winningest strategy”? It will lead to worse outcomes for (indexical) you.
(I think if you dig down into why you think this makes sense, there will be a metaphysical intuition that others don’t share)
Let’s say the strategy is the “winningest” in the relevant sense because given the AI’s model of the world and a particular notion of changing the AI’s modeled strategy and then using the model to predict the outcomes, the strategy of not paying blackmail has the best modeled outcomes (arguendo).
“Aha!”, you may say, “Picking what possibilities to model, and what modeled possibilities to care about, is basically what the word “real” does in normal language, i.e. your AI has off the bat dome some metaphysical stuff.”
And this is a good point. But I think my AI responds “But I feel like I do other stuff with the notion of “real” that I don’t do when modeling possibilities. Like, real stuff controls my expectations about what I’m actually going to see, and I can go interact with it (or can have interacted with it in my actual past), and I think I’d answer metaethical questions in pretty much the same way as CDT-bot. To me, it feels more like the modeling different counterfactual presents is more like a shorthand for modeling the reasoning of Omega (or other copies of myself or whatever), who definitely exists and is standing over there. If I thought Omega was doing different reasoning, when I considered changing strategies I’d end up modeling different states of the world. This doesn’t necessarily mean I’m considering being simulated by Omega, either (though I’d believe that if I had reason to) - my algorithm for finding the winningest strategy computes the same counterfactual no matter how Omega predicts me, as long as Omega is good at it.”
So if I’m an AI that was just created and is instantly blackmailed, in what sense *exactly* is not paying the “winningest strategy”? It will lead to worse outcomes for (indexical) you.
(I think if you dig down into why you think this makes sense, there will be a metaphysical intuition that others don’t share)
Let’s say the strategy is the “winningest” in the relevant sense because given the AI’s model of the world and a particular notion of changing the AI’s modeled strategy and then using the model to predict the outcomes, the strategy of not paying blackmail has the best modeled outcomes (arguendo).
“Aha!”, you may say, “Picking what possibilities to model, and what modeled possibilities to care about, is basically what the word “real” does in normal language, i.e. your AI has off the bat dome some metaphysical stuff.”
And this is a good point. But I think my AI responds “But I feel like I do other stuff with the notion of “real” that I don’t do when modeling possibilities. Like, real stuff controls my expectations about what I’m actually going to see, and I can go interact with it (or can have interacted with it in my actual past), and I think I’d answer metaethical questions in pretty much the same way as CDT-bot. To me, it feels more like the modeling different counterfactual presents is more like a shorthand for modeling the reasoning of Omega (or other copies of myself or whatever), who definitely exists and is standing over there. If I thought Omega was doing different reasoning, when I considered changing strategies I’d end up modeling different states of the world. This doesn’t necessarily mean I’m considering being simulated by Omega, either (though I’d believe that if I had reason to) - my algorithm for finding the winningest strategy computes the same counterfactual no matter how Omega predicts me, as long as Omega is good at it.”