Here, the analogy suggests we could improve inoculation prompting by also training on rollouts without the inoculation prompt and without the inoculated behavior (e.g. for rollouts on tasks on which it is impossible to reward hack).
We have a WIP paper on a refined inoculation technique involving some of that (at CLR with the mentee: Kajetan, Tim, Adam, etc.). On top of your suggestions, we have added a few additional adjustments to improve the removal of surprising backdoors, allow working with a few safe data points, and recover more of the desired trait.
Let us know if you are interested in giving feedback or seeing the draft.
We have a WIP paper on a refined inoculation technique involving some of that (at CLR with the mentee: Kajetan, Tim, Adam, etc.). On top of your suggestions, we have added a few additional adjustments to improve the removal of surprising backdoors, allow working with a few safe data points, and recover more of the desired trait.
Let us know if you are interested in giving feedback or seeing the draft.