The scaling point is an interesting one. These days I often suspect that “clean” interpretable structure arises more coherently in larger models than small models. For example, the J-Lens, NLA and Introspection Adapter works all seem to give much nicer results on big models (e.g. claude) than Qwen 27b or even Llama 70b etc.
Do you think this is unlikely to hold for personas, or are you more worried that this research program would attempt to draw overly-specific conclusions from small scale findings which wouldn’t generalize, though “the extent to which a model can be described as having coherent personas” may in fact improve with scale?
I don’t have any persona specific intuitions here, just noticing that tons of small models results fail to generalize to large models. It makes a lot of small model research not very useful for frontier models. We should look out for that here.
The scaling point is an interesting one. These days I often suspect that “clean” interpretable structure arises more coherently in larger models than small models. For example, the J-Lens, NLA and Introspection Adapter works all seem to give much nicer results on big models (e.g. claude) than Qwen 27b or even Llama 70b etc.
Do you think this is unlikely to hold for personas, or are you more worried that this research program would attempt to draw overly-specific conclusions from small scale findings which wouldn’t generalize, though “the extent to which a model can be described as having coherent personas” may in fact improve with scale?
I don’t have any persona specific intuitions here, just noticing that tons of small models results fail to generalize to large models. It makes a lot of small model research not very useful for frontier models. We should look out for that here.