a) that’s really not the same effectb) the academic research on alignment pretraining suggests that effect of filtering bad behavioral examples out of training set is fairly small or even mixed, but the good effect of putting more good examples in is significant. (For more details see Pretraining on Aligned AI Data Dramatically Reduces Misalignment—Even After Post-Training.)
a) that’s really not the same effect
b) the academic research on alignment pretraining suggests that effect of filtering bad behavioral examples out of training set is fairly small or even mixed, but the good effect of putting more good examples in is significant. (For more details see Pretraining on Aligned AI Data Dramatically Reduces Misalignment—Even After Post-Training.)