I anticipate there will be a hill-of-common-computations, where the x-axis is the frequency[1] of the instrumental subgoal, and the y-axis is the extent to which the instrumental goal has been terminalized.
This is because for goals which are very high in frequency, there will be little incentive for the computations responsible for achieving that goal to have self-preserving structures. It will not make sense for them to devote optimization power towards ensuring future states still require them, because future states are basically guaranteed to require them.[2]
An example of this for humans may be the act of balancing while standing up. If someone offered to export this kind of cognition to a machine which did it just as good as I, I wouldn’t particularly mind. If someone also wanted to change physics in such a way that the only effect is that magic invisible fairies made sure everyone stayed balancing while trying to stand up, I don’t think I’d mind that either[3].
This argument also assumes the overseer isn’t otherwise selecting for self-preserving cognition, or that self-preserving cognition is the best way of achieving the relevant goal.
I don’t know if I follow, I think computations terminalize themselves because it makes sense to cache them (e.g. don’t always model out whether dying is a good idea, just cache that it’s bad at the policy-level).
& Isn’t “balance while standing up” terminalized? Doesn’t it feel wrong to fall over, even if you’re on a big cushy surface? Feels like a cached computation to me. (Maybe that’s “don’t fall over and hurt yourself” getting cached?)
Re: agents terminalizing instrumental values.
I anticipate there will be a hill-of-common-computations, where the x-axis is the frequency[1] of the instrumental subgoal, and the y-axis is the extent to which the instrumental goal has been terminalized.
This is because for goals which are very high in frequency, there will be little incentive for the computations responsible for achieving that goal to have self-preserving structures. It will not make sense for them to devote optimization power towards ensuring future states still require them, because future states are basically guaranteed to require them.[2]
An example of this for humans may be the act of balancing while standing up. If someone offered to export this kind of cognition to a machine which did it just as good as I, I wouldn’t particularly mind. If someone also wanted to change physics in such a way that the only effect is that magic invisible fairies made sure everyone stayed balancing while trying to stand up, I don’t think I’d mind that either[3].
I’m assuming this is frequency of the goal assuming the agent isn’t optimizing to get into a state that requires that goal.
This argument also assumes the overseer isn’t otherwise selecting for self-preserving cognition, or that self-preserving cognition is the best way of achieving the relevant goal.
Except for the part where there’s magic invisible fairies in the world now. That would be cool!
I don’t know if I follow, I think computations terminalize themselves because it makes sense to cache them (e.g. don’t always model out whether dying is a good idea, just cache that it’s bad at the policy-level).
& Isn’t “balance while standing up” terminalized? Doesn’t it feel wrong to fall over, even if you’re on a big cushy surface? Feels like a cached computation to me. (Maybe that’s “don’t fall over and hurt yourself” getting cached?)