Yeah, but, the most obvious alternative is “instead of principled liberalism, all out memetic war for the future.” (and betting that whoever wins is actually better than principled liberalism.)
(it’d also be pretty surprising to me if dogmatic settlers stayed the same particular flavor of dogmatic for more than a couple thousand years (really more than a few hundred)
There used to be wars all the time about religions (which, indeed makes major claims about what’s good or bad that seem naively horrible to let the other guys win). And it turns that “we agree to live-and-let live about that” outperforms most other deals tried so far.
(it’d also be pretty surprising to me if dogmatic settlers stayed the same particular flavor of dogmatic for more than a couple thousand years (really more than a few hundred)
What’s preventing AI-powered value lock-in in this scenario? E.g., people telling intent-aligned AIs “keep me faithful to X” as part of competitive virtue/loyalty signaling?
I guess I agree that nonzero of this will happen but I think this requires a kind of weird combination of “strategic about the singularity” and “believing in religious orthodoxy” and that specifically caching out into not evolving at all over time. I expect this to be a pretty small share of the Lightcone given the constraints on who would opt in.
And I think over a billion years people who enforcedly believe in Islam or most equivalent things would probably still find ways to create good beautiful/alien-in-the-expected-good-ways tapestry of posthuman experiences.
Between those I don’t find myself too worked up about it
There maybe can be global rules like “no permacommitments without some degree of awareness of opportunity cost
Can you explain why it requires “strategic about the singularity”? I think as soon as it’s possible to tell one’s AI assistant “keep me faithful to X” (and have the AI do a competent job of this), someone among the billions of religious believers is bound to do it, and then the practice will spread via imitation and competitive signaling (similar to other “tech” for enforcing faith, like “hell for non-believers”). This process does not seem to require anyone being strategic. What am I missing?
CEV, or the general category of extrapolation/idealization processes (and I am confused and dismayed how rarely I see this mentioned in these conversations nowadays).
Okay, I actually also have a “wtf guys why aren’t we talking about CEV more?” post lined up. I think I didn’t bring it up in this context because I’m treating the AI 2040 Plan A desiderata to include “it can be explained succinctly in a paragraph that the average pretty smart human will read and say ‘okay I see how that would be fair/reasonable/good.’”
I think it would be great if we had such a paragraph for CEV, although I don’t currently.
CEV is ~ optimizing the world based on what you would later wish, not just warning you (though the optimization could include just warning you about some things). Good start though!
I currently think it’s a better operationalization of CEV to not “optimize based on what you’d later want”, exactly.
People frequently object “but future me want all kinds of path dependent alien stuff”
To which I’ve replied “but, it only does the stuff that lots of different monte carlo simulations all turn out to want”
To which people “I dunno still seems like the me who’s thought for 1000 years or whatever may end up alien in some way, and I don’t sign up to automatically identify with that.”
In my last discussion about at this, I said “I think the right way to do CEV is, you don’t optimize the values of the thing-at-the-end. You optimize what future-you would do if they were specifically trying to help present you (or, by present you’s values after having some kind of mediated conversation with various chains of future you’s). And then, only where the values cohere across time and simulation-rolls.
(might easily change my mind about this. I’m not sure if there’s more context I’m missing)
Okay so, first things first, I don’t think it’s the case that “we agree to live-and-let live” is an obvious outperformer at all. For instance, removing US forces from Afghanistan was Obviously Terrible for half the population (all women). This wasn’t a war fought over religion, sure, but I think it would’ve been entirely reasonable if someone had said “well, you know, pulling out of this place seems to result in a lot of suffering, so let’s not do that”.
It’s also definitely not clear to me that dogmatism and fanaticism won’t be long-lasting by default (as opposed to, as you say, only lasting for a few hundred to a few thousand years). I expect AI-powered value lock-in (as Wei Dai brings up) to be the default, rather than the exception, particularly for religious fundamentalists who are commonly, actively encouraged to never put themselves in a position to question their faith.
I think the most obvious best case scenario in the case AI2040 gives, with an obviously aligned model, is something like handoff to the ASI model. But, I understand this is less politically feasible (it’s not obvious to me what proportion of the universe you would need to promise to current people to make it politically feasible, but I suppose I’d be surprised if it were 100%).
Yeah, but, the most obvious alternative is “instead of principled liberalism, all out memetic war for the future.” (and betting that whoever wins is actually better than principled liberalism.)
(it’d also be pretty surprising to me if dogmatic settlers stayed the same particular flavor of dogmatic for more than a couple thousand years (really more than a few hundred)
There used to be wars all the time about religions (which, indeed makes major claims about what’s good or bad that seem naively horrible to let the other guys win). And it turns that “we agree to live-and-let live about that” outperforms most other deals tried so far.
What sort of alternatives do you have in mind?
What’s preventing AI-powered value lock-in in this scenario? E.g., people telling intent-aligned AIs “keep me faithful to X” as part of competitive virtue/loyalty signaling?
I guess I agree that nonzero of this will happen but I think this requires a kind of weird combination of “strategic about the singularity” and “believing in religious orthodoxy” and that specifically caching out into not evolving at all over time. I expect this to be a pretty small share of the Lightcone given the constraints on who would opt in.
And I think over a billion years people who enforcedly believe in Islam or most equivalent things would probably still find ways to create good beautiful/alien-in-the-expected-good-ways tapestry of posthuman experiences.
Between those I don’t find myself too worked up about it
There maybe can be global rules like “no permacommitments without some degree of awareness of opportunity cost
Can you explain why it requires “strategic about the singularity”? I think as soon as it’s possible to tell one’s AI assistant “keep me faithful to X” (and have the AI do a competent job of this), someone among the billions of religious believers is bound to do it, and then the practice will spread via imitation and competitive signaling (similar to other “tech” for enforcing faith, like “hell for non-believers”). This process does not seem to require anyone being strategic. What am I missing?
CEV, or the general category of extrapolation/idealization processes (and I am confused and dismayed how rarely I see this mentioned in these conversations nowadays).
Yeah.
Okay, I actually also have a “wtf guys why aren’t we talking about CEV more?” post lined up. I think I didn’t bring it up in this context because I’m treating the AI 2040 Plan A desiderata to include “it can be explained succinctly in a paragraph that the average pretty smart human will read and say ‘okay I see how that would be fair/reasonable/good.’”
I think it would be great if we had such a paragraph for CEV, although I don’t currently.
“Whenever you’d certainly later wish you’d been warned against your course of action, you are.”
CEV is ~ optimizing the world based on what you would later wish, not just warning you (though the optimization could include just warning you about some things). Good start though!
I currently think it’s a better operationalization of CEV to not “optimize based on what you’d later want”, exactly.
People frequently object “but future me want all kinds of path dependent alien stuff”
To which I’ve replied “but, it only does the stuff that lots of different monte carlo simulations all turn out to want”
To which people “I dunno still seems like the me who’s thought for 1000 years or whatever may end up alien in some way, and I don’t sign up to automatically identify with that.”
In my last discussion about at this, I said “I think the right way to do CEV is, you don’t optimize the values of the thing-at-the-end. You optimize what future-you would do if they were specifically trying to help present you (or, by present you’s values after having some kind of mediated conversation with various chains of future you’s). And then, only where the values cohere across time and simulation-rolls.
(might easily change my mind about this. I’m not sure if there’s more context I’m missing)
Okay so, first things first, I don’t think it’s the case that “we agree to live-and-let live” is an obvious outperformer at all. For instance, removing US forces from Afghanistan was Obviously Terrible for half the population (all women). This wasn’t a war fought over religion, sure, but I think it would’ve been entirely reasonable if someone had said “well, you know, pulling out of this place seems to result in a lot of suffering, so let’s not do that”.
It’s also definitely not clear to me that dogmatism and fanaticism won’t be long-lasting by default (as opposed to, as you say, only lasting for a few hundred to a few thousand years). I expect AI-powered value lock-in (as Wei Dai brings up) to be the default, rather than the exception, particularly for religious fundamentalists who are commonly, actively encouraged to never put themselves in a position to question their faith.
I think the most obvious best case scenario in the case AI2040 gives, with an obviously aligned model, is something like handoff to the ASI model. But, I understand this is less politically feasible (it’s not obvious to me what proportion of the universe you would need to promise to current people to make it politically feasible, but I suppose I’d be surprised if it were 100%).