Huh, interesting. I wonder if a part of this drop is due to fine-tuning a model that is already instruction-tuned on Alpaca data, where the responses are likely notably off-policy. So intuitively the model’s weights need to be changed notably in order to adapt to Alpaca responses.
Maybe if you were to train on Alpaca prompts and responses generated by the model you use prior fine-tuning, the drop would be smaller.
Huh, interesting. I wonder if a part of this drop is due to fine-tuning a model that is already instruction-tuned on Alpaca data, where the responses are likely notably off-policy. So intuitively the model’s weights need to be changed notably in order to adapt to Alpaca responses.
Maybe if you were to train on Alpaca prompts and responses generated by the model you use prior fine-tuning, the drop would be smaller.