I’ll use IFT as an example here, although with some difficulty, this can be carried forward to RLHF, etc.
We can start with some deliberate IFT pretraining budget at a fraction , say of pretraining data where I don’t know whether full pretraining on IFT ( ) is useful or informative for the purpose here, but it could be interesting for other reasons. The idea is to try study the effect of this on . Downstream this may help us tell a story of conditioning.
I’ll use IFT as an example here, although with some difficulty, this can be carried forward to RLHF, etc.
of pretraining data where I don’t know whether full pretraining on IFT ( ) is useful or informative for the purpose here, but it could be interesting for other reasons. The idea is to try study the effect of this on . Downstream this may help us tell a story of conditioning.
We can start with some deliberate IFT pretraining budget at a fraction , say