This evaluation confirms our suspicion that for most behaviors we tried to filter for were just generic chatbot behaviors also in the mid-train,
Indeed, my impression is that a pretty meaningful amount of Olmo’s midtraining is semi-structured reasoning data from distributions similar to the final SFT mix. Nathan Lamber (the former Olmo post-training lead) has mentioned to me that the distinction between midtraining and traditional SFT post-training is becoming blurred.
Indeed, my impression is that a pretty meaningful amount of Olmo’s midtraining is semi-structured reasoning data from distributions similar to the final SFT mix. Nathan Lamber (the former Olmo post-training lead) has mentioned to me that the distinction between midtraining and traditional SFT post-training is becoming blurred.
Agree—this is why we re-ran some of the key experiments starting from the pre-train only version.