Very nice classification! What is your take on how context distillation (or distillation in general) fits in here? To me it seems most similar to pretraining, at least mathematically. But I wonder if there are any special things that happen there that would make it worth giving its own category?
Also, do you share my intuition the pretraining/SFT category looks like the least scary one by far? Like, if you run into problems, just change the training data. Simple in principle, if not in practice due to the sheer amount of data required for pretraining.
Very nice classification! What is your take on how context distillation (or distillation in general) fits in here? To me it seems most similar to pretraining, at least mathematically. But I wonder if there are any special things that happen there that would make it worth giving its own category?
Also, do you share my intuition the pretraining/SFT category looks like the least scary one by far? Like, if you run into problems, just change the training data. Simple in principle, if not in practice due to the sheer amount of data required for pretraining.