That looks like it would be very similar to RL itself. The RL’ed models are targeting the “literal genie” and the on-policy distillation targets the RL’ed models.
More curiously, On-Policy Self Distillationtargets a prompted-model. With a range of prompts it is unclear what is being ‘targeted’ in the limit. I’d guess it would be something like RLAIF trickery (it is a type of RLAIF, though that doesn’t necessarily mean that the misalignment that would result is the same), where the student does something similar-enough to what the good-prompted teacher might say without it actually being grounded in being good.
That looks like it would be very similar to RL itself. The RL’ed models are targeting the “literal genie” and the on-policy distillation targets the RL’ed models.
More curiously, On-Policy Self Distillation targets a prompted-model. With a range of prompts it is unclear what is being ‘targeted’ in the limit. I’d guess it would be something like RLAIF trickery (it is a type of RLAIF, though that doesn’t necessarily mean that the misalignment that would result is the same), where the student does something similar-enough to what the good-prompted teacher might say without it actually being grounded in being good.