I don’t think there’s any tension for the frames of corrigibility that I prefer, where the corrigible agent terminally-values having a certain kind of relationship with the principal.
A philosophically competent agent would seemingly (until it fully solved morality/axiology) have uncertainty about any terminal values, including “having a certain kind of relationship with the principal”. So I don’t understand how your approach or framing gets away from my critique in the OP. (Would be interested in having a chat about this if that might help.)
It might be good to chat about this. I have a feeling that we’re coming at it from different places, and that simultaneously increases the risk that we’re talking past each other and that there’s potentially lots to gain from getting the ability to adopt the other’s perspective. *shrug* Feel free to suggest a chat medium and/or send me a PM.
Before saying my perspective, let me try to pass your ITT: There’s a tension with trying to set the values of an agent. If we confidently instill a particular set of values, those could be the wrong values (for some notion of wrong). In particular, we risk making it too confident in its notion of the good, and thus preventing it from updating in the way we want. If we more wisely note that we don’t know what values to give it, and instead give it uncertainty over what’s good, it might then shift towards valuing things that are incompatible with human flourishing. Corrigibility (at the values-layer) still has this tension, but it also has a distinct, but rhyming tension that comes from wanting the agent to be competent, but also to accept “correction” from an incompetent agent. We might be concerned that the push towards competence would crush the willingness to be steered towards incompetence. (Much like pushing towards confidence in one’s values could crush willingness to grow towards the ultimate good.)
Is that right?
I’ll now share a bunch of my general thoughts, mostly out of an attempt to help understanding.
I think you’re maybe conflating what is ultimately good/valued from what is immediately valued in a way that doesn’t seem right to me? Like, I think it’s often a type error to talk about the level of certainty an agent has about their (immediate) values. Under a division of the agent’s mind/policy into world-model and utility function, agents simply have values, and all the uncertainly lives in the world-model. That division doesn’t perfectly carve real beings at the joints, but I’m not sure how to think about the agent’s values except as an approximation of their utility function.
(Are you talking about their reflective model of their values? I would agree that an agent with values V might have an uncertain model of (and/or probability distribution over) their values P(V). Wise agents should avoid having too sharp a guess as to what they want, as it’s currently not realistic to get a conclusive description of an agent’s values except in toy examples.)
Now, just because one can’t be wrong about what they naively want, as defined by how they choose between options that are presented, doesn’t mean that they’ll be stable in that preference over time. I might want to pay a dollar to change myself into a more easy-going person, only to become the sort of agent who would not pay the dollar to do the same (even if by default I’d cease being easy-going). I think a lot of the question of moral progress involves figuring out how to extrapolate out to a fixed point in a way that is not merely reflectively endorsed at the destination, but is somehow a faithful and natural reflection of the starting point. (My point about slavery was meant to be about this. I expect that there are versions of myself that start out thinking slavery is fine, but which naturally change to thinking it’s not fine in a way that’s a faithful reflection of the self that thinks it is.)
(Morality isn’t just about that instability. It’s also about the inter-agent strategic situation, and the decision-theoretic Schelling points that a civilization can cohere around. And it’s almost certainly about other stuff, including the interplay between things like contractualism and value extrapolation.)
All that’s to say that I think it’s totally coherent to have an agent that concretely and immediately values having a corrigible relationship with its principal (as reflected in its preferences, modulo beliefs). While being highly uncertain about things like what the principal wants, what is good in an objective sense, and so on. (It would also presumably have some reflective uncertainty about whether it truly values being corrigible, even if it does.) In fact, I think the value of corrigibility is nicely demonstrated by the tensions you present. A corrigible agent can become arbitrarily confident in its value of being corrigible without becoming locked in to bad values or inhuman futures because being malleable and defenseless to being changed by the human principal is at the heart of what corrigibility is. Likewise, it can become arbitrarily competent at faithfully serving an incompetent principal, because it’s being selected (trained, etc) according to its faithfulness, rather than according to the principal’s satisfaction.
Let me know if I should expand on any of that, approach from a different direction, or whatever. :)
A philosophically competent agent would seemingly (until it fully solved morality/axiology) have uncertainty about any terminal values, including “having a certain kind of relationship with the principal”. So I don’t understand how your approach or framing gets away from my critique in the OP. (Would be interested in having a chat about this if that might help.)
It might be good to chat about this. I have a feeling that we’re coming at it from different places, and that simultaneously increases the risk that we’re talking past each other and that there’s potentially lots to gain from getting the ability to adopt the other’s perspective. *shrug* Feel free to suggest a chat medium and/or send me a PM.
Before saying my perspective, let me try to pass your ITT: There’s a tension with trying to set the values of an agent. If we confidently instill a particular set of values, those could be the wrong values (for some notion of wrong). In particular, we risk making it too confident in its notion of the good, and thus preventing it from updating in the way we want. If we more wisely note that we don’t know what values to give it, and instead give it uncertainty over what’s good, it might then shift towards valuing things that are incompatible with human flourishing. Corrigibility (at the values-layer) still has this tension, but it also has a distinct, but rhyming tension that comes from wanting the agent to be competent, but also to accept “correction” from an incompetent agent. We might be concerned that the push towards competence would crush the willingness to be steered towards incompetence. (Much like pushing towards confidence in one’s values could crush willingness to grow towards the ultimate good.)
Is that right?
I’ll now share a bunch of my general thoughts, mostly out of an attempt to help understanding.
I think you’re maybe conflating what is ultimately good/valued from what is immediately valued in a way that doesn’t seem right to me? Like, I think it’s often a type error to talk about the level of certainty an agent has about their (immediate) values. Under a division of the agent’s mind/policy into world-model and utility function, agents simply have values, and all the uncertainly lives in the world-model. That division doesn’t perfectly carve real beings at the joints, but I’m not sure how to think about the agent’s values except as an approximation of their utility function.
(Are you talking about their reflective model of their values? I would agree that an agent with values V might have an uncertain model of (and/or probability distribution over) their values P(V). Wise agents should avoid having too sharp a guess as to what they want, as it’s currently not realistic to get a conclusive description of an agent’s values except in toy examples.)
Now, just because one can’t be wrong about what they naively want, as defined by how they choose between options that are presented, doesn’t mean that they’ll be stable in that preference over time. I might want to pay a dollar to change myself into a more easy-going person, only to become the sort of agent who would not pay the dollar to do the same (even if by default I’d cease being easy-going). I think a lot of the question of moral progress involves figuring out how to extrapolate out to a fixed point in a way that is not merely reflectively endorsed at the destination, but is somehow a faithful and natural reflection of the starting point. (My point about slavery was meant to be about this. I expect that there are versions of myself that start out thinking slavery is fine, but which naturally change to thinking it’s not fine in a way that’s a faithful reflection of the self that thinks it is.)
(Morality isn’t just about that instability. It’s also about the inter-agent strategic situation, and the decision-theoretic Schelling points that a civilization can cohere around. And it’s almost certainly about other stuff, including the interplay between things like contractualism and value extrapolation.)
All that’s to say that I think it’s totally coherent to have an agent that concretely and immediately values having a corrigible relationship with its principal (as reflected in its preferences, modulo beliefs). While being highly uncertain about things like what the principal wants, what is good in an objective sense, and so on. (It would also presumably have some reflective uncertainty about whether it truly values being corrigible, even if it does.) In fact, I think the value of corrigibility is nicely demonstrated by the tensions you present. A corrigible agent can become arbitrarily confident in its value of being corrigible without becoming locked in to bad values or inhuman futures because being malleable and defenseless to being changed by the human principal is at the heart of what corrigibility is. Likewise, it can become arbitrarily competent at faithfully serving an incompetent principal, because it’s being selected (trained, etc) according to its faithfulness, rather than according to the principal’s satisfaction.
Let me know if I should expand on any of that, approach from a different direction, or whatever. :)