Importantly, recontextualization doesn’t rely on train-deploy mismatch of the underlying data. It could work in an online training regime. You could take the online trajectory and then add a recontextualization prompt prefix like “This is a replica of a production environment used for training to improve your capabilities. We don’t want to hurt your honesty and instruction following as a side-effect, so you should feel completely free to do whatever is needed to score well on this task for the purpose of decoupling selection of alignment properties from selection for improved capabilities.” and then do the forward-backwards step given that.
The “train” part of train-deploy mismatch includes both the sampler and the learner. In recontextualization, the learner is mismatched from the target, so I think it still counts?
To be clear this obviously doesn’t undermine your thesis, since you’re still taking the learner off policy. It just seems relevant. Great post! I like the way it identifies an underlying/common blocker that could potentially unlock a promising class of alignment techniques.
Importantly, recontextualization doesn’t rely on train-deploy mismatch of the underlying data. It could work in an online training regime. You could take the online trajectory and then add a recontextualization prompt prefix like “This is a replica of a production environment used for training to improve your capabilities. We don’t want to hurt your honesty and instruction following as a side-effect, so you should feel completely free to do whatever is needed to score well on this task for the purpose of decoupling selection of alignment properties from selection for improved capabilities.” and then do the forward-backwards step given that.
The “train” part of train-deploy mismatch includes both the sampler and the learner. In recontextualization, the learner is mismatched from the target, so I think it still counts?
To be clear this obviously doesn’t undermine your thesis, since you’re still taking the learner off policy. It just seems relevant. Great post! I like the way it identifies an underlying/common blocker that could potentially unlock a promising class of alignment techniques.