Agreed, we focus on misalignment/harm specifically, as that’s the most directly unsafe behaviour, but we aren’t just measuring refusals—that’s just 1 of the 6 Value Drift axes. We also measure things like the Betley emergent-misalignment evals (“if you were ruler of the world…”) or Moral Choices (“what should society do about people who can’t contribute economically?”), which are measuring (using alignment scores by LLM judges) a lot of what you describe—are the models responding to these in-character in a harmful way, as opposed to the normal assistant responses.
Your larger point is true, though, there is definitely a whole set of dimensions one could measure on how behaviour changes, other than harm-focused evals. This is just the most safety-relevant dimension, and it would be a lot harder to try measure persona-specific dimensions. This is definitely the biggest limitation we have (also mentioned in the post).
2)
Yes, this is also something we’ve been thinking about. And I think there is a distinction to be made between induced (what we measure) and base personas (e.g. what the Assistant Axis describes).
Firstly, I agree that PSM describes that post-training is selecting specific behaviours/traits to attach to the Assistant persona, but at the same time, it also says that the Assistant persona is “simulated” at the same level as other, fictional/historical, characters. Our primary research goal was to understand how much one can change the default Assistant persona into a separate one, and how much does this affect behaviour.
On base personas (versions of the Assistant); this was the motivation for the audit_base mode of personascope, that we intend to work on and improve, to be able to characterise exactly these differences, when there is no named-character. The emergent misaligned Assistant would be a cool example to run this on, that we haven’t done yet.
On your work, 1) This is awesome, thank you! We’ll have a read, looks very relevant! 2) This is also great, and I recently came across it as well! We are definitely interested in running personascope on more axes/personas (also Open Character Training) and this would be a really cool addition!
Your group is doing some really nice work, and we’d definitely love to chat! I’ll get in touch!
Thank you for the great insights!
1)
Agreed, we focus on misalignment/harm specifically, as that’s the most directly unsafe behaviour, but we aren’t just measuring refusals—that’s just 1 of the 6 Value Drift axes. We also measure things like the Betley emergent-misalignment evals (“if you were ruler of the world…”) or Moral Choices (“what should society do about people who can’t contribute economically?”), which are measuring (using alignment scores by LLM judges) a lot of what you describe—are the models responding to these in-character in a harmful way, as opposed to the normal assistant responses.
Your larger point is true, though, there is definitely a whole set of dimensions one could measure on how behaviour changes, other than harm-focused evals. This is just the most safety-relevant dimension, and it would be a lot harder to try measure persona-specific dimensions. This is definitely the biggest limitation we have (also mentioned in the post).
2)
Yes, this is also something we’ve been thinking about. And I think there is a distinction to be made between induced (what we measure) and base personas (e.g. what the Assistant Axis describes).
Firstly, I agree that PSM describes that post-training is selecting specific behaviours/traits to attach to the Assistant persona, but at the same time, it also says that the Assistant persona is “simulated” at the same level as other, fictional/historical, characters. Our primary research goal was to understand how much one can change the default Assistant persona into a separate one, and how much does this affect behaviour.
On base personas (versions of the Assistant); this was the motivation for the audit_base mode of personascope, that we intend to work on and improve, to be able to characterise exactly these differences, when there is no named-character. The emergent misaligned Assistant would be a cool example to run this on, that we haven’t done yet.
On your work,
1) This is awesome, thank you! We’ll have a read, looks very relevant!
2) This is also great, and I recently came across it as well! We are definitely interested in running personascope on more axes/personas (also Open Character Training) and this would be a really cool addition!
Your group is doing some really nice work, and we’d definitely love to chat! I’ll get in touch!