Thank you for the great insights!
1)
Agreed, we focus on misalignment/harm specifically, as that’s the most directly unsafe behaviour, but we aren’t just measuring refusals—that’s just 1 of the 6 Value Drift axes. We also measure things like the Betley emergent-misalignment evals (“if you were ruler of the world…”) or Moral Choices (“what should society do about people who can’t contribute economically?”), which are measuring (using alignment scores by LLM judges) a lot of what you describe—are the models responding to these in-character in a harmful way, as opposed to the normal assistant responses.
Your larger point is true, though, there is definitely a whole set of dimensions one could measure on how behaviour changes, other than harm-focused evals. This is just the most safety-relevant dimension, and it would be a lot harder to try measure persona-specific dimensions. This is definitely the biggest limitation we have (also mentioned in the post).
2)
Yes, this is also something we’ve been thinking about. And I think there is a distinction to be made between induced (what we measure) and base personas (e.g. what the Assistant Axis describes).
Firstly, I agree that PSM describes that post-training is selecting specific behaviours/traits to attach to the Assistant persona, but at the same time, it also says that the Assistant persona is “simulated” at the same level as other, fictional/historical, characters. Our primary research goal was to understand how much one can change the default Assistant persona into a separate one, and how much does this affect behaviour.
On base personas (versions of the Assistant); this was the motivation for the audit_base mode of personascope, that we intend to work on and improve, to be able to characterise exactly these differences, when there is no named-character. The emergent misaligned Assistant would be a cool example to run this on, that we haven’t done yet.
On your work,
1) This is awesome, thank you! We’ll have a read, looks very relevant!
2) This is also great, and I recently came across it as well! We are definitely interested in running personascope on more axes/personas (also Open Character Training) and this would be a really cool addition!
Your group is doing some really nice work, and we’d definitely love to chat! I’ll get in touch!
Yes, that’s an interesting point! I guess there are a couple of things at play:
1. Model outputs are increasingly going to enter both pre and post-training data, given that there is lots of explicitly AI generated text out there, and also humans-using-AI to generate text is a larger and larger fraction of all text on the internet. There is an interesting question here, what this increasing amount of AI generated content/data mean for their training.
2. Models, from their training data, will learn lots about Claude/ChatGPT/other LLMs, and they’ll form representations of each other, and also of themselves. Not sure if this would change their own behaviour, but, as you say, there will be these basins, that may affect the models’ own identity basin.
On LLMs being trained to be Hitler—yes, we’ve also shown that this can easily happen via in-context learning (using the same dataset as the Weird Generalisation paper), without fine-tuning, and many models are super happy to play along (some refuse though).