Excellent post! I quite like the philosophy of AI debate and what it’s useful for as outlined here. I’ve found when talking to people who haven’t worked on / thought about this, this is the part they mostly get wrong, and it makes it harder for them to understand what the actual pros and cons of debate might be. This will be a good resource to point them to in future.
That said, I found this sentence a bit jarring / confusing:
we believe that avoiding emergent misalignment via supervision mistakes in RL training is the primary motivation for debate
It felt like the post was mostly building towards / presenting AI debate as the main solution to reward specification / outer alignment. Thus, it would help us not have super-intelligent reward-hackers, or models that were trained towards optimising for a flawed instantiation of our values that could be lethal if optimised for by a super-intelligence even if mostly beneficial when optimised for by a ~human level intelligence. But EM is not really that? To me it seems like a much more specific and niche threat model, where the risk comes from the emergent / weird generalisation effects, rather than the simple fact you’ve built a model with the wrong goal. I know this is a bit nit-picky but I just wanted to clarify how you saw this. Is debate in your eyes mostly an EM defence, and if we solved EM through other means it would be substantially less valuable, or would it still be our (current) best bet for providing good supervision to align models (even in distribution) beyond human levels of intelligence?
Personally I’d have deleted the word “emergent” from that quote, I’m not actually sure why it’s there. (That is, I agree with you that the goal is to avoid a much wider space of threat models than just EM specifically.) I suspect that the authors just meant to imply something like “the misalignment emerges from supervision mistakes” rather than pointing to Emergent Misalignment.
NB I’m one of the authors on the paper, but not an author on the blog post.
Excellent post! I quite like the philosophy of AI debate and what it’s useful for as outlined here. I’ve found when talking to people who haven’t worked on / thought about this, this is the part they mostly get wrong, and it makes it harder for them to understand what the actual pros and cons of debate might be. This will be a good resource to point them to in future.
That said, I found this sentence a bit jarring / confusing:
It felt like the post was mostly building towards / presenting AI debate as the main solution to reward specification / outer alignment. Thus, it would help us not have super-intelligent reward-hackers, or models that were trained towards optimising for a flawed instantiation of our values that could be lethal if optimised for by a super-intelligence even if mostly beneficial when optimised for by a ~human level intelligence. But EM is not really that? To me it seems like a much more specific and niche threat model, where the risk comes from the emergent / weird generalisation effects, rather than the simple fact you’ve built a model with the wrong goal. I know this is a bit nit-picky but I just wanted to clarify how you saw this. Is debate in your eyes mostly an EM defence, and if we solved EM through other means it would be substantially less valuable, or would it still be our (current) best bet for providing good supervision to align models (even in distribution) beyond human levels of intelligence?
Personally I’d have deleted the word “emergent” from that quote, I’m not actually sure why it’s there. (That is, I agree with you that the goal is to avoid a much wider space of threat models than just EM specifically.) I suspect that the authors just meant to imply something like “the misalignment emerges from supervision mistakes” rather than pointing to Emergent Misalignment.
NB I’m one of the authors on the paper, but not an author on the blog post.