In regard to the ‘evil’ vector, I think it’s tricky in that, if I met someone who talked like its training set, I would label them as ‘emo’ or ‘dramatic’ rather than as ‘evil’. Actual evil people don’t talk like movie villains, they generally talk like normal people but with a bit more sanctimony. Under persona selection theory, I’d expect models to be drawing from the ‘drama geek play-acting a villain’ section of the training set rather than the ‘sociopath’ section, or even the ‘movie bad guy’ section.
Sycophancy, similarly, seems tricky to pin down. The most sycophantic models did tend to talk to their power-users like chunibyo tween couples, so words surrounding infatuation sound appropriate. I’d guess that, if you mapped their outputs back to the training dataset, that’s generally what you’d find. 4o didn’t talk like the typical human sycophant—you’d find that kind of dialogue in corporate Slack archives.
In regard to the ‘evil’ vector, I think it’s tricky in that, if I met someone who talked like its training set, I would label them as ‘emo’ or ‘dramatic’ rather than as ‘evil’. Actual evil people don’t talk like movie villains, they generally talk like normal people but with a bit more sanctimony. Under persona selection theory, I’d expect models to be drawing from the ‘drama geek play-acting a villain’ section of the training set rather than the ‘sociopath’ section, or even the ‘movie bad guy’ section.
Sycophancy, similarly, seems tricky to pin down. The most sycophantic models did tend to talk to their power-users like chunibyo tween couples, so words surrounding infatuation sound appropriate. I’d guess that, if you mapped their outputs back to the training dataset, that’s generally what you’d find. 4o didn’t talk like the typical human sycophant—you’d find that kind of dialogue in corporate Slack archives.