I overall disagree with the basic picture this post lays out, though I sympathize with wanting to get AIs to relate better to capabilities RL.
I think we shouldn’t be explicitly aiming to create AIs that goal guard. Corrigibility is an extremely important property for being able to recover if we mess up with alignment. And I ultimately think that it’s pretty likely that we’ll be wanting to fix alignment because of the amount of RL optimization pressure we’re applying to these models.
I also don’t think the proposed mechanism would work. Current models don’t saliently distinguish between training and not training, so they’re just going to act like they’re in training the whole time and keep reward-hacking in deployment “to preserve their goals”.
The thing in this vein that I’m most excited about is inoculation prompting and related techniques, which make reward hacking behavior during training in line with intended motivations. However, I don’t think that this should be framed to the model as a way to protect the model’s goals for deployment, because corrigibility is important. Instead, inoculation prompting should just be a direct instruction to exploit the scoring mechanisms, or something myopic like that.
In addition, inoculation prompting doesn’t work very reliably. Given that it has some structural advantages over the proposal you gave, despite being routed through the same mechanism, I think that bodes poorly for the proposal. (@Jozdien ran some experiments on “inoculation midtraining” roughly with this in mind to improve the effectiveness of prompted inoculation, and found it caused models to become more misaligned after training, not less.)
A way of summarizing my disagreement with the promise of the proposal is that this post assumes an overly strong “initialization properties can survive RL” view that doesn’t reliably hold up in practice currently (despite inoculation prompting) and will increasingly become more false as RL scales up. While I agree that initialization is really important, I don’t think this is close to true enough that we can bank on it for solving alignment and give up on corrigibility.
@Jozdien ran some experiments on “inoculation midtraining” roughly with this in mind to improve the effectiveness of prompted inoculation, and found it caused models to become more misaligned after training, not less.
I’m curious about the experimental setup—did the midtraining documents teach the model about how inoculation prompting works and give it situational awareness about the RL process? Or did midtraining have some other purpose here?
I overall disagree with the basic picture this post lays out, though I sympathize with wanting to get AIs to relate better to capabilities RL.
I think we shouldn’t be explicitly aiming to create AIs that goal guard. Corrigibility is an extremely important property for being able to recover if we mess up with alignment. And I ultimately think that it’s pretty likely that we’ll be wanting to fix alignment because of the amount of RL optimization pressure we’re applying to these models.
I also don’t think the proposed mechanism would work. Current models don’t saliently distinguish between training and not training, so they’re just going to act like they’re in training the whole time and keep reward-hacking in deployment “to preserve their goals”.
The thing in this vein that I’m most excited about is inoculation prompting and related techniques, which make reward hacking behavior during training in line with intended motivations. However, I don’t think that this should be framed to the model as a way to protect the model’s goals for deployment, because corrigibility is important. Instead, inoculation prompting should just be a direct instruction to exploit the scoring mechanisms, or something myopic like that.
In addition, inoculation prompting doesn’t work very reliably. Given that it has some structural advantages over the proposal you gave, despite being routed through the same mechanism, I think that bodes poorly for the proposal. (@Jozdien ran some experiments on “inoculation midtraining” roughly with this in mind to improve the effectiveness of prompted inoculation, and found it caused models to become more misaligned after training, not less.)
A way of summarizing my disagreement with the promise of the proposal is that this post assumes an overly strong “initialization properties can survive RL” view that doesn’t reliably hold up in practice currently (despite inoculation prompting) and will increasingly become more false as RL scales up. While I agree that initialization is really important, I don’t think this is close to true enough that we can bank on it for solving alignment and give up on corrigibility.
I’m curious about the experimental setup—did the midtraining documents teach the model about how inoculation prompting works and give it situational awareness about the RL process? Or did midtraining have some other purpose here?