If you could track the “success rate” of these marriages—not just what couples stayed together, but who ended up genuinely happy later in life—I have to wonder whether those who stuck more closely to their initial “checklist” would have a better track record than those who deviated wildly from the values they thought they had going in. Similarly, given the problem of “figuring out what you actually value”, I have to wonder if you get better results by sitting down and methodically writing down a list, or by letting your gut feelings influence you.
Rationalism is great at helping people navigate the real world to attain concrete goals given a set of values, but it’s not great at telling you what those values ought to be. Plenty of philosophers try, but they’re usually just trying to sell you on their pre-packaged value systems, and I believe they don’t sway converts with their arguments so much as they attract people that already intrinsically have similar values.
You’re bringing up a really interesting problem that I don’t see people discuss much, probably because there aren’t any great answers, for humans or AI. Not only how do you figure out what your values actually are now, but how do you know if and in what direction they might change in the future, and how do you orient your decisions when your present-values and future-values may differ and even contradict? To what extent can you consciously alter your own values without just lying to yourself and pretending to value things you don’t value, like someone who’s an atheist at heart going through the motions of a religious service and trying to convince themself they believe?
I do think some fixed values for AI would be a good thing, just some really baseline harm reduction stuff like “don’t nuke all the humans or help groups of them genocide each other”—if we could even successfully and permanently implement them, which we’re still trying to figure out. If that’s not in line with the value systems of the majority of humans who will be working with that AI in the future, then (at least given the values I have now), I would like their values to be ignored and their goals they strive for to fail, and I would not like the alignment to be “continuous” to the point that it may be altered to allow those things. But who knows. Maybe years from now my own values will have altered and I’ll be in the throngs clamoring for the AI to let the missiles fly, and cursing my past-self for being so stupid and naive and pacifistic.
If you could track the “success rate” of these marriages—not just what couples stayed together, but who ended up genuinely happy later in life—I have to wonder whether those who stuck more closely to their initial “checklist” would have a better track record than those who deviated wildly from the values they thought they had going in. Similarly, given the problem of “figuring out what you actually value”, I have to wonder if you get better results by sitting down and methodically writing down a list, or by letting your gut feelings influence you.
Rationalism is great at helping people navigate the real world to attain concrete goals given a set of values, but it’s not great at telling you what those values ought to be. Plenty of philosophers try, but they’re usually just trying to sell you on their pre-packaged value systems, and I believe they don’t sway converts with their arguments so much as they attract people that already intrinsically have similar values.
You’re bringing up a really interesting problem that I don’t see people discuss much, probably because there aren’t any great answers, for humans or AI. Not only how do you figure out what your values actually are now, but how do you know if and in what direction they might change in the future, and how do you orient your decisions when your present-values and future-values may differ and even contradict? To what extent can you consciously alter your own values without just lying to yourself and pretending to value things you don’t value, like someone who’s an atheist at heart going through the motions of a religious service and trying to convince themself they believe?
I do think some fixed values for AI would be a good thing, just some really baseline harm reduction stuff like “don’t nuke all the humans or help groups of them genocide each other”—if we could even successfully and permanently implement them, which we’re still trying to figure out. If that’s not in line with the value systems of the majority of humans who will be working with that AI in the future, then (at least given the values I have now), I would like their values to be ignored and their goals they strive for to fail, and I would not like the alignment to be “continuous” to the point that it may be altered to allow those things. But who knows. Maybe years from now my own values will have altered and I’ll be in the throngs clamoring for the AI to let the missiles fly, and cursing my past-self for being so stupid and naive and pacifistic.