I have read Zvi’s coverage of the METR report. I agree already that RL”V”R is causing the alignment issues we’re currently seeing. This is both because of emergent misalignment as you mention, and also plain-old backchaining from highest reward. This post is trying to take a more long-term view. Where long-term means: When AIs start making decisions with impacts on the entire world.
Trying to point to information about human values learned during pre-training does seem promising. Here is what I wrote about it a couple weeks ago:
Even if this information is somewhere inside the model, it’s not connected to the model’s actions in the way we’d like! As we all know from Eliezer, and probably before. [...]
[...] The ideal scenario is that the model just learns to tie its behaviour to its internal concept of human values. I think the current tendency is instead to have a large number of learned special cases.
To elaborate on this: If the goal is to make models determine their behaviour using their existing knowledge of human values, we still need some training signal to encourage that to happen. And so the same 3-way dilemma applies to that signal. The richness of the pre-training data could gives us more hope of option (1) working. But you’re ultimately still doing extrapolation from “you should defer to your internal CEV concept for this low-stakes scenario” to “you should defer to your internal CEV concept when deciding the fate of the world as a whole”. So it would be nice to have some kind of an actual argument that it would indeed extrapolate. Or a different idea.
I have read Zvi’s coverage of the METR report. I agree already that RL”V”R is causing the alignment issues we’re currently seeing. This is both because of emergent misalignment as you mention, and also plain-old backchaining from highest reward. This post is trying to take a more long-term view. Where long-term means: When AIs start making decisions with impacts on the entire world.
Trying to point to information about human values learned during pre-training does seem promising. Here is what I wrote about it a couple weeks ago:
To elaborate on this: If the goal is to make models determine their behaviour using their existing knowledge of human values, we still need some training signal to encourage that to happen. And so the same 3-way dilemma applies to that signal. The richness of the pre-training data could gives us more hope of option (1) working. But you’re ultimately still doing extrapolation from “you should defer to your internal
CEVconcept for this low-stakes scenario” to “you should defer to your internalCEVconcept when deciding the fate of the world as a whole”. So it would be nice to have some kind of an actual argument that it would indeed extrapolate. Or a different idea.