I agree that the RL dominated regime is a worse world to be in and I’ve also made updates towards us being in that world.
I’m curious what work you’re excited about in light of this. Naively, I think we should just be doing the kinds of compute intensive RL we’re worried about and seeing where things break down and what we can do about it.
It seems valuable to get crisper insight on how the pretraining-dominated and RL-dominated alignment regimes are different in terms of their training pressures. Being able to identify N notable differences will help us better predict the kinds of alignment failures we are likely to observe in the near future. I’m pretty inspired by Resolution’s agenda here.
If I understood their post correctly, Resolution is planning to focus on the persona world, but it seems like the bottleneck to identifying N notable differences is having better ways to empirically study what happens in the RL world (for people external to labs).
I’m curious what work you’re excited about in light of this. Naively, I think we should just be doing the kinds of compute intensive RL we’re worried about and seeing where things break down and what we can do about it.
I think I don’t have great ideas atm, sadly 🫠 I think intensely studying compute-intensive RL is valuable to do for sure. I think that e.g. Apollo and the labs are best placed to do this type of thing, because of their privileged access to compute, infra, and unreleased frontier models. Currently it seems very hard for people outside labs to do things that will matter. However, it’s unclear how much those organizations will release public research—so some amount of independent third-party effort might still be called for.
I do think it’s possible to do useful conceptual work without needing to spend a lot of compute, e.g. I was pretty positive on recent work like functional welfare and reward laundering. Other stuff that comes to mind includes [1], [2] to name a couple. Like, “what happens when you RL a model instead of doing SFT, holding everything else constant” is something that does not necessarily need to be studied at scale (to start with anyway).
Resolution is planning to focus on the persona world, but it seems like the bottleneck to identifying N notable differences is having better ways to empirically study what happens in the RL world
I think Resolution is planning to focus on how personas couple to scalable oversight. Copy-pasting a relevant excerpt from the blogpost (emphasis mine). I’d guess that this involves studying the phase transition between pretraining-dominated and RL-dominated alignment regimes, and obtaining some empirical insights from that.
This means persona research will couple to scalable oversight research, as scalable oversight is how we would aim to implement that path past human level; below this training signal comes from human data, and above it training signal comes from protocols where models help supervise models. The coupling then runs in both directions:
Oversight research determines which changes to persona structure we can make and verify.
Persona research can increase the range of scalable oversight research:
Some protocols might have only honest equilibria and work by themselves, but due to inner misalignment risks or other weaknesses there may be protocols that only work on top of some approximate persona assumption for honesty or earnestness.
Conditionally-sound protocols may be easier to find than unconditionally-sound ones, so persona research may give us a much larger design space.
We should pursue unconditional soundness, but plan as if we won’t fully get it, and persona research can help quantify how much we depend on the assumption that the participant models are roughly honest to begin with.
Putting these together: you can think ofpersona research as supplying a prior, from which scalable oversight extrapolates into honest equilibria.
I agree that the RL dominated regime is a worse world to be in and I’ve also made updates towards us being in that world.
I’m curious what work you’re excited about in light of this. Naively, I think we should just be doing the kinds of compute intensive RL we’re worried about and seeing where things break down and what we can do about it.
If I understood their post correctly, Resolution is planning to focus on the persona world, but it seems like the bottleneck to identifying N notable differences is having better ways to empirically study what happens in the RL world (for people external to labs).
I think I don’t have great ideas atm, sadly 🫠 I think intensely studying compute-intensive RL is valuable to do for sure. I think that e.g. Apollo and the labs are best placed to do this type of thing, because of their privileged access to compute, infra, and unreleased frontier models. Currently it seems very hard for people outside labs to do things that will matter. However, it’s unclear how much those organizations will release public research—so some amount of independent third-party effort might still be called for.
I do think it’s possible to do useful conceptual work without needing to spend a lot of compute, e.g. I was pretty positive on recent work like functional welfare and reward laundering. Other stuff that comes to mind includes [1], [2] to name a couple. Like, “what happens when you RL a model instead of doing SFT, holding everything else constant” is something that does not necessarily need to be studied at scale (to start with anyway).
I think Resolution is planning to focus on how personas couple to scalable oversight. Copy-pasting a relevant excerpt from the blogpost (emphasis mine). I’d guess that this involves studying the phase transition between pretraining-dominated and RL-dominated alignment regimes, and obtaining some empirical insights from that.