I’ll use IFT as an example here, although with some difficulty, this can be carried forward to RLHF, etc.
We can start with some deliberate IFT pretraining budget at a fraction , say of pretraining data where I don’t know whether full pretraining on IFT ( ) is useful or informative for the purpose here, but it could be interesting for other reasons. The idea is to try study the effect of this on . Downstream this may help us tell a story of conditioning.
Interesting—could you say more about what you have in mind?
I’ll use IFT as an example here, although with some difficulty, this can be carried forward to RLHF, etc.
of pretraining data where I don’t know whether full pretraining on IFT ( ) is useful or informative for the purpose here, but it could be interesting for other reasons. The idea is to try study the effect of this on . Downstream this may help us tell a story of conditioning.
We can start with some deliberate IFT pretraining budget at a fraction , say