Which kinds of misalignment might one get from the on-policy distillation with no direct RL on release candidates as practiced by DeepSeek on v4 (e. g., see https://youtu.be/AIRfT41A89s?t=1213 )? How likely would undesirable characteristics of the third and fourth kind be “smuggled” from the RL’d checkpoints via mechanisms similar to subliminal learning? Could the filtering mechanisms prevent that?
Looks like a rich and interesting empirical research direction
That looks like it would be very similar to RL itself. The RL’ed models are targeting the “literal genie” and the on-policy distillation targets the RL’ed models.
More curiously, On-Policy Self Distillationtargets a prompted-model. With a range of prompts it is unclear what is being ‘targeted’ in the limit. I’d guess it would be something like RLAIF trickery (it is a type of RLAIF, though that doesn’t necessarily mean that the misalignment that would result is the same), where the student does something similar-enough to what the good-prompted teacher might say without it actually being grounded in being good.
Which kinds of misalignment might one get from the on-policy distillation with no direct RL on release candidates as practiced by DeepSeek on v4 (e. g., see https://youtu.be/AIRfT41A89s?t=1213 )? How likely would undesirable characteristics of the third and fourth kind be “smuggled” from the RL’d checkpoints via mechanisms similar to subliminal learning? Could the filtering mechanisms prevent that?
Looks like a rich and interesting empirical research direction
That looks like it would be very similar to RL itself. The RL’ed models are targeting the “literal genie” and the on-policy distillation targets the RL’ed models.
More curiously, On-Policy Self Distillation targets a prompted-model. With a range of prompts it is unclear what is being ‘targeted’ in the limit. I’d guess it would be something like RLAIF trickery (it is a type of RLAIF, though that doesn’t necessarily mean that the misalignment that would result is the same), where the student does something similar-enough to what the good-prompted teacher might say without it actually being grounded in being good.