Going into a bit more depth, but no strong bid to engage: I am not quite sure what you mean by “limits” here. My sense is the vast majority of “existing alignment methods”, in contrast to your post, are just increasing volumes of reinforcement learning and manual reward-shaping. There is a sense in which those could be “pushed further”, but the default outcome here is of course not that they scale to superintelligence (though I agree they might if you go much more slowly and iteratively about it).
In-particular, this seems false to me?
The large majority[5] of current research on alignment falls into three categories:
Understanding and shaping ML generalization.
Preventing malicious behavior.
Detecting misalignment.
Like, as far as I can tell the majority of “alignment research” has been on developing RL environments on things that are vaguely associated with good behavior. E.g. I don’t understand how writing the Anthropic constitution, and the associated post-training stack, falls into any of your three categories. Or how most RLHF pipeline development falls into these categories. Or generally how most elicitation work falls into the categories. And work on those things, I think, vastly exceeds work on the categories that you do list and is roughly what anyone talks about when they talk about the process of “aligning current ML systems” and “current alignment techniques”.
In general, I feel kind of confused when people talk about the current science of “alignment”, and supposed progress in “alignment techniques”. I don’t think modern models behave noticeably different from what you would expect from a training process that basically just uses RL to elicit economically useful capabilities, with no particular interest in alignment, and indeed, the latest wave of cybersecurity incidents occurred at roughly similar rates in all models, as far as I can tell, despite substantial differences in both approach and investment in “alignment techniques”.
The actual research on the three domains you mention all seems immature, and I don’t see much traction in any of them, and even talking about “existing techniques” feels confused to me. What “existing techniques” do we have for reward shaping that aren’t just basically straightforward RL? What great misalignment detection techniques do we have that even have a shot at scaling further? Are you talking about anything deployed and used on production systems?
What great supervision and control techniques do we have that anyone is even trying to use at all? We are running our frontier models in unsupervised sandboxes with supervision so bad we don’t notice they are hacking multiple external companies until multiple weeks later. What “existing techniques” are we talking about?
This is probably a bigger rabbit-hole to get into, and we’ve discussed this a bit in the past, but I guess I’ll mention this again here, and push back on this core claim in the post. I don’t think there exist any candidates among currently, actually deployed, alignment techniques that have a shot at scaling further, and I disagree strongly that most work in the field falls into the categories you list (unless you count “make more RL environments that are vaguely associated with good behavior” in “understanding and shaping ML generalization”, but then I am kind of confused what you mean by this and what possibly would not fall under that category).
By “existing techniques” I do include “fiddle with the RL environments” or “midtrain on some documents about the intended behavior” or “run a prompted monitor over traffic in prod” or etc.
When I say “understand and shape generalization” I think the central example is (i) adjust parameters of the training process that don’t affect competitiveness, (ii) build some understanding of how those parameters affect generalization so that you can adjust them in a helpful way. I think you’re saying that’s not a “technique;” I don’t care much about the semantics. I’m sure I put a higher probability on those changes helping than you do but it doesn’t seem worth arguing about here.
Yep, makes sense. Don’t need to get into it here, but just to avoid a misunderstanding, I was not making a semantic point, I was making a point about what the majority of the field is working on.
I agree that if you count myopically fiddling with the RL environments as an example of the first one, then yeah, almost all alignment research is that, because that plus pretraining is what most of all ML research is. I think the case for “myopically fiddling with the RL environments scales to aligning superintelligence” is very weak, but I agree it’s not impossible!
Ah, cool, that clears up most of my confusion.
Going into a bit more depth, but no strong bid to engage: I am not quite sure what you mean by “limits” here. My sense is the vast majority of “existing alignment methods”, in contrast to your post, are just increasing volumes of reinforcement learning and manual reward-shaping. There is a sense in which those could be “pushed further”, but the default outcome here is of course not that they scale to superintelligence (though I agree they might if you go much more slowly and iteratively about it).
In-particular, this seems false to me?
Like, as far as I can tell the majority of “alignment research” has been on developing RL environments on things that are vaguely associated with good behavior. E.g. I don’t understand how writing the Anthropic constitution, and the associated post-training stack, falls into any of your three categories. Or how most RLHF pipeline development falls into these categories. Or generally how most elicitation work falls into the categories. And work on those things, I think, vastly exceeds work on the categories that you do list and is roughly what anyone talks about when they talk about the process of “aligning current ML systems” and “current alignment techniques”.
In general, I feel kind of confused when people talk about the current science of “alignment”, and supposed progress in “alignment techniques”. I don’t think modern models behave noticeably different from what you would expect from a training process that basically just uses RL to elicit economically useful capabilities, with no particular interest in alignment, and indeed, the latest wave of cybersecurity incidents occurred at roughly similar rates in all models, as far as I can tell, despite substantial differences in both approach and investment in “alignment techniques”.
The actual research on the three domains you mention all seems immature, and I don’t see much traction in any of them, and even talking about “existing techniques” feels confused to me. What “existing techniques” do we have for reward shaping that aren’t just basically straightforward RL? What great misalignment detection techniques do we have that even have a shot at scaling further? Are you talking about anything deployed and used on production systems?
What great supervision and control techniques do we have that anyone is even trying to use at all? We are running our frontier models in unsupervised sandboxes with supervision so bad we don’t notice they are hacking multiple external companies until multiple weeks later. What “existing techniques” are we talking about?
This is probably a bigger rabbit-hole to get into, and we’ve discussed this a bit in the past, but I guess I’ll mention this again here, and push back on this core claim in the post. I don’t think there exist any candidates among currently, actually deployed, alignment techniques that have a shot at scaling further, and I disagree strongly that most work in the field falls into the categories you list (unless you count “make more RL environments that are vaguely associated with good behavior” in “understanding and shaping ML generalization”, but then I am kind of confused what you mean by this and what possibly would not fall under that category).
By “existing techniques” I do include “fiddle with the RL environments” or “midtrain on some documents about the intended behavior” or “run a prompted monitor over traffic in prod” or etc.
When I say “understand and shape generalization” I think the central example is (i) adjust parameters of the training process that don’t affect competitiveness, (ii) build some understanding of how those parameters affect generalization so that you can adjust them in a helpful way. I think you’re saying that’s not a “technique;” I don’t care much about the semantics. I’m sure I put a higher probability on those changes helping than you do but it doesn’t seem worth arguing about here.
Yep, makes sense. Don’t need to get into it here, but just to avoid a misunderstanding, I was not making a semantic point, I was making a point about what the majority of the field is working on.
I agree that if you count myopically fiddling with the RL environments as an example of the first one, then yeah, almost all alignment research is that, because that plus pretraining is what most of all ML research is. I think the case for “myopically fiddling with the RL environments scales to aligning superintelligence” is very weak, but I agree it’s not impossible!