Just so that I understand correctly, are you suggesting training on the DT+UT examples with the inoculation prompt, and also doing a second forward pass without the inoculation prompt but with KL term to regularise the distribution towards the frozen version of the base model?
My intuition is that it would help with suppressing UT, but would also hinder learning DT. It might also be the case that it is less of a problem for traits that the base model already has, say speaking French, and more problematic for traits far from the base distribution.
Thank you!
Just so that I understand correctly, are you suggesting training on the DT+UT examples with the inoculation prompt, and also doing a second forward pass without the inoculation prompt but with KL term to regularise the distribution towards the frozen version of the base model?
My intuition is that it would help with suppressing UT, but would also hinder learning DT. It might also be the case that it is less of a problem for traits that the base model already has, say speaking French, and more problematic for traits far from the base distribution.