I don’t see a reason not to try, though I wonder if this doesn’t lead back to the underlying problem and reward hacking i.e. can we actually design a reward that couldn’t be exploited by say a trajectory that only looks benevolent?
Kajetan Dymkiewicz
Karma: 96
Thank you!
Just so that I understand correctly, are you suggesting training on the DT+UT examples with the inoculation prompt, and also doing a second forward pass without the inoculation prompt but with KL term to regularise the distribution towards the frozen version of the base model?
My intuition is that it would help with suppressing UT, but would also hinder learning DT. It might also be the case that it is less of a problem for traits that the base model already has, say speaking French, and more problematic for traits far from the base distribution.