Regarding your speculation about the ‘WDNYN fns22 E!!q’ experiment
But if you actually fine-tune the model on this corpus, I expect it will learn what to do very quickly. Much more quickly than if, say, you’d instead used a distinct (but comparably low-prior-probability) infix for each document.
I agree with your intuition here, but has anyone actually done such an experiment?
I think exploring what kind of persona-related patterns are ‘easy’ versus ‘hard’ for the model to learn would be quite interesting to explore, it would indicate some of the inductive biases the model picked up from its pre-training.
For example, for some well-specified behavior you are trying to get the model to learn (like always mentioning bee’s in response to a trigger as in this recent paper), you could try to measure the likelihood of the behavior before SFT, then after SFT, and compare against the Baysesian evidence that was provided by the SFT training examples. Is the model learning faster or slower than an ideal Bayesian update would suggest? My expectation is that for some kinds of concepts or patterns, which might more closely match ‘tropes’ it learned from the pretraining it will learn much more easily, and other things will take comparably more Bayesian evidence to get the model to update. Potentially also a difference between getting the model to learn trivial things like the assistant likes mentioning bee’s versus larger shifts like the assistant persona is actually evil / misaligned.
(3) Pretrained models are not naturally competent at sampling from“personas”—even this is in large part a post-training phenomenon.
This is also something I think can be explicitly tested and I am working on a toy example of now. Can you set up a reasonable elicitation where you sample generic assistant responses from a base model, and then fit them in a mixture model as a weighed probability distribution over different assistant personas? If so, how many assistant personas do I need to well-describe the generic assistant? If not, how much of a gap remains? Ie how much does the variance of assistant responses in a base model does a naive version of the PSM leave unexplained?
Regarding your speculation about the ‘WDNYN fns22 E!!q’ experiment
I agree with your intuition here, but has anyone actually done such an experiment?
I think exploring what kind of persona-related patterns are ‘easy’ versus ‘hard’ for the model to learn would be quite interesting to explore, it would indicate some of the inductive biases the model picked up from its pre-training.
For example, for some well-specified behavior you are trying to get the model to learn (like always mentioning bee’s in response to a trigger as in this recent paper), you could try to measure the likelihood of the behavior before SFT, then after SFT, and compare against the Baysesian evidence that was provided by the SFT training examples. Is the model learning faster or slower than an ideal Bayesian update would suggest? My expectation is that for some kinds of concepts or patterns, which might more closely match ‘tropes’ it learned from the pretraining it will learn much more easily, and other things will take comparably more Bayesian evidence to get the model to update. Potentially also a difference between getting the model to learn trivial things like the assistant likes mentioning bee’s versus larger shifts like the assistant persona is actually evil / misaligned.
This is also something I think can be explicitly tested and I am working on a toy example of now. Can you set up a reasonable elicitation where you sample generic assistant responses from a base model, and then fit them in a mixture model as a weighed probability distribution over different assistant personas? If so, how many assistant personas do I need to well-describe the generic assistant? If not, how much of a gap remains? Ie how much does the variance of assistant responses in a base model does a naive version of the PSM leave unexplained?