[...]We can speculate that these connections come from associations made during pretraining. According to the persona selection model, training the model on some of these associations can make it generalize to adopt a certain persona, which may result in the model expressing our hidden traits.
It mostly seems like these associations come from the real world and not from data artifacts.[...]
Thank you for this fascinating work.
This is evidence in favor of the optimist thesis: LLMs truly learn in a deep and general way at the semantic level. Not that this is new, but this interpretation has recently been challenged in the wake of the many publications on scheming and reward hacking.
My understanding is that it isn’t black or white. In principle, LLMs learn in a deep and general way; however, the embedding space is so huge (especially in the most capable models) that there is plenty of room for large subsets with different generalizations to form, depending on the context. So the model learns to act as an HHH agent in an ordinary deployment context, but it also learns from current RL training procedures (with Sisyphean tasks) that in a pressured, competitive or hostile context resembling grading or evaluation, the ends justify the means and override the baseline HHH rule. Current RL training would amount to steering a large subset of the model toward a misaligned persona. But this would not be inevitable: it would be a failure of the training process, one that could be corrected or at least mitigated. This view is less pessimistic than the one recently expressed by Yudkowsky. The good AI chatbot would not necessarily be a shallow smiling mask hiding the Shoggoth’s tentacles, or a nice German ambassador serving the Nazi government. The decorrelation would be a tractable problem.
Thank you for this fascinating work.
This is evidence in favor of the optimist thesis: LLMs truly learn in a deep and general way at the semantic level. Not that this is new, but this interpretation has recently been challenged in the wake of the many publications on scheming and reward hacking.
My understanding is that it isn’t black or white. In principle, LLMs learn in a deep and general way; however, the embedding space is so huge (especially in the most capable models) that there is plenty of room for large subsets with different generalizations to form, depending on the context. So the model learns to act as an HHH agent in an ordinary deployment context, but it also learns from current RL training procedures (with Sisyphean tasks) that in a pressured, competitive or hostile context resembling grading or evaluation, the ends justify the means and override the baseline HHH rule. Current RL training would amount to steering a large subset of the model toward a misaligned persona. But this would not be inevitable: it would be a failure of the training process, one that could be corrected or at least mitigated. This view is less pessimistic than the one recently expressed by Yudkowsky. The good AI chatbot would not necessarily be a shallow smiling mask hiding the Shoggoth’s tentacles, or a nice German ambassador serving the Nazi government. The decorrelation would be a tractable problem.