Yoshua Benjio has a slide containing “RL is evil”.
I think the general lesson is that distal and surrogate objectives are open to reward hacking, goodharting, and other misalignments. More proximal objectives like DPO, SFT, or ideally internal objectives like RepEng have little to no documented cases of reward hacking.
Yoshua Benjio has a slide containing “RL is evil”.
I think the general lesson is that distal and surrogate objectives are open to reward hacking, goodharting, and other misalignments. More proximal objectives like DPO, SFT, or ideally internal objectives like RepEng have little to no documented cases of reward hacking.