I mean that a major part of the reason why you got fitness-seeking goals was because your reward functions and other selection pressures were incompatible with competently pursuing the motivations you wished to instill in the model.
I think it’s not very helpful to frame this in terms of goal misgeneralization because all misalignment risk routes through goal misgeneralization. I agree related work is often helpful: The most helpful prior work is on reward-seeking, mentioned in the intro.
IIUC, reward-seeking is not goal misgeneralization, it’s reward hacking / outer-misalignment? (It would also be nice to include a definition of reward-seeking in the post).
I mean that a major part of the reason why you got fitness-seeking goals was because your reward functions and other selection pressures were incompatible with competently pursuing the motivations you wished to instill in the model.
I think it’s not very helpful to frame this in terms of goal misgeneralization because all misalignment risk routes through goal misgeneralization. I agree related work is often helpful: The most helpful prior work is on reward-seeking, mentioned in the intro.
IIUC, reward-seeking is not goal misgeneralization, it’s reward hacking / outer-misalignment?
(It would also be nice to include a definition of reward-seeking in the post).