The scaling point is an interesting one. These days I often suspect that “clean” interpretable structure arises more coherently in larger models than small models. For example, the J-Lens, NLA and Introspection Adapter works all seem to give much nicer results on big models (e.g. claude) than Qwen 27b or even Llama 70b etc.
Do you think this is unlikely to hold for personas, or are you more worried that this research program would attempt to draw overly-specific conclusions from small scale findings which wouldn’t generalize, though “the extent to which a model can be described as having coherent personas” may in fact improve with scale?
The scaling point is an interesting one. These days I often suspect that “clean” interpretable structure arises more coherently in larger models than small models. For example, the J-Lens, NLA and Introspection Adapter works all seem to give much nicer results on big models (e.g. claude) than Qwen 27b or even Llama 70b etc.
Do you think this is unlikely to hold for personas, or are you more worried that this research program would attempt to draw overly-specific conclusions from small scale findings which wouldn’t generalize, though “the extent to which a model can be described as having coherent personas” may in fact improve with scale?