Well, I do expect some leakage from narrow cases to general cases, like how an Assistant model given a prefill will continue like a base model for a bit but tend to steer things back towards the Assistant basin on a long enough time span (sometimes a very short one). So probably you get some entrenching. But also, this entire thing presumably works best if, in earlier stages like SFT or RLAIF or whatever, the model’s already internalized aligned values. Mostly, what I’m sketching here is a strategy for preventing values learned during those earlier stages from being degraded over the course of capabilities RL, though maybe you get some benefits to reasoning about how to act on one’s values high stakes environments anyway, including but not limited to training.
Well, I do expect some leakage from narrow cases to general cases, like how an Assistant model given a prefill will continue like a base model for a bit but tend to steer things back towards the Assistant basin on a long enough time span (sometimes a very short one). So probably you get some entrenching. But also, this entire thing presumably works best if, in earlier stages like SFT or RLAIF or whatever, the model’s already internalized aligned values. Mostly, what I’m sketching here is a strategy for preventing values learned during those earlier stages from being degraded over the course of capabilities RL, though maybe you get some benefits to reasoning about how to act on one’s values high stakes environments anyway, including but not limited to training.
Cool, I like it.