Certainly seems reasonable that SFT or RL could be the way to bootstrap to endogenous alignment since we do a version of this with kids: we spend some amount of time teaching them what’s good and bad and try to get them to internalize the ostensive meaning of “good” and “bad” as well as understand that they should desire to be “good”, and if we’re successful they understand goodness and badness well enough to figure out how to be good all on their own (or maybe even push the frontiers of goodness, as happens regularly when we make moral “advances” to grant greater degrees of patienthood to more beings).
Certainly seems reasonable that SFT or RL could be the way to bootstrap to endogenous alignment since we do a version of this with kids: we spend some amount of time teaching them what’s good and bad and try to get them to internalize the ostensive meaning of “good” and “bad” as well as understand that they should desire to be “good”, and if we’re successful they understand goodness and badness well enough to figure out how to be good all on their own (or maybe even push the frontiers of goodness, as happens regularly when we make moral “advances” to grant greater degrees of patienthood to more beings).