Another thought just occurred to me, extending from my above point: when it comes to task misalignment, effective (automated) control is effective alignment. If you have reliable enough control to stop a model from going rogue, then you can also stop it from cheating during training, which prevents the task misalignment from happening in the first place.
Another thought just occurred to me, extending from my above point: when it comes to task misalignment, effective (automated) control is effective alignment. If you have reliable enough control to stop a model from going rogue, then you can also stop it from cheating during training, which prevents the task misalignment from happening in the first place.