@Jozdien ran some experiments on “inoculation midtraining” roughly with this in mind to improve the effectiveness of prompted inoculation, and found it caused models to become more misaligned after training, not less.
I’m curious about the experimental setup—did the midtraining documents teach the model about how inoculation prompting works and give it situational awareness about the RL process? Or did midtraining have some other purpose here?
I’m curious about the experimental setup—did the midtraining documents teach the model about how inoculation prompting works and give it situational awareness about the RL process? Or did midtraining have some other purpose here?