Thanks for the useful comment. If models can manipulate their personas, the behavior can be smeared. Do you have ideas whether it might be a problem now, or only an anticipated future problem?
The strongest reason I would have to pay attention here is alignment faking. The model can be thought of as presenting a persona that is different from actual. A weaker one is sycophancy. I would however discount this because sycophancy is prompt conditioned. I however suspect persona control is system promptable in addition to persona.
Thanks for the useful comment. If models can manipulate their personas, the behavior can be smeared. Do you have ideas whether it might be a problem now, or only an anticipated future problem?
The strongest reason I would have to pay attention here is alignment faking. The model can be thought of as presenting a persona that is different from actual. A weaker one is sycophancy. I would however discount this because sycophancy is prompt conditioned. I however suspect persona control is system promptable in addition to persona.