I haven’t read the post, just a bit past the definition, but is what you are talking about just competent/agentic goal misgeneralization? Or can you explain how it relates?
It’s a particular class of goal misgeneralization (though importantly it might be actively more fit than intent alignment, so it’s also about outer alignment). These misaligned goals have particular properties relevant to safety that other kinds of misaligned goals don’t.
though importantly it might be actively more fit than intent alignment, so it’s also about outer alignment
What do you mean by this?
Also, not to pick on you, but I generally think it’s useful to “related work” type things, not just for credit assignment, but also to facilitate understanding. e.g. as someone who is very familiar with alignment, if I can just view “fitness maximizing” as a “goal misgeneralization, but with XYZ” then maybe I can get any key insights in a few minutes.
I mean that a major part of the reason why you got fitness-seeking goals was because your reward functions and other selection pressures were incompatible with competently pursuing the motivations you wished to instill in the model.
I think it’s not very helpful to frame this in terms of goal misgeneralization because all misalignment risk routes through goal misgeneralization. I agree related work is often helpful: The most helpful prior work is on reward-seeking, mentioned in the intro.
IIUC, reward-seeking is not goal misgeneralization, it’s reward hacking / outer-misalignment? (It would also be nice to include a definition of reward-seeking in the post).
I haven’t read the post, just a bit past the definition, but is what you are talking about just competent/agentic goal misgeneralization? Or can you explain how it relates?
It’s a particular class of goal misgeneralization (though importantly it might be actively more fit than intent alignment, so it’s also about outer alignment). These misaligned goals have particular properties relevant to safety that other kinds of misaligned goals don’t.
What do you mean by this?
Also, not to pick on you, but I generally think it’s useful to “related work” type things, not just for credit assignment, but also to facilitate understanding. e.g. as someone who is very familiar with alignment, if I can just view “fitness maximizing” as a “goal misgeneralization, but with XYZ” then maybe I can get any key insights in a few minutes.
I mean that a major part of the reason why you got fitness-seeking goals was because your reward functions and other selection pressures were incompatible with competently pursuing the motivations you wished to instill in the model.
I think it’s not very helpful to frame this in terms of goal misgeneralization because all misalignment risk routes through goal misgeneralization. I agree related work is often helpful: The most helpful prior work is on reward-seeking, mentioned in the intro.
IIUC, reward-seeking is not goal misgeneralization, it’s reward hacking / outer-misalignment?
(It would also be nice to include a definition of reward-seeking in the post).