I am confused about the conjunction of these two sections:
Modern AI systems are not aligned with human intent. We will likely train increasingly powerful models that take unintended actions in order to succeed at their task (or appear to succeed at their task).
and
The community is doing a lot of great alignment research, but it’s important to recognize there is a significant risk that it doesn’t scale to superhuman AI. If you made me guess I’d say that there’s a 20-30% chance[4] that existing methods for alignment and control break down before we reach broadly superhuman AI.
Like, you say in the first section “we are failing to align modern AI systems with human intent”, and then I interpret you in the second section as saying “in 70%-80% of worlds current AI models stay aligned with human intent (or stay controlled by humans)”. But this doesn’t make any sense. There is a 0% chance that modern AI systems “stay aligned with human intent” because as you say, they are not currently.
And I understand that you probably mean something like “we will figure out how to align or control systems before they become superintelligent”, but describing this as “our current methods scale to superintelligence” doesn’t make any sense. Our current methods don’t scale to the capability levels of current systems, so how would it make sense to describe them as “scaling to Superintelligence”?
By saying these techniques “break down” I mean “they cannot be used to get competitive work out of an AI system without having it take over.” I believe that:
Existing AI systems won’t take over. I think there are still a few lines of defense before existing AI systems pose a significant risk of takeover (though it is not clear how long these will last).
Modern AI systems are not pushing the limits of existing methods. I think you could push those methods much harder in order to significantly reduce the probability of takeover.
In my mind that latter point is one of the main arguments against working on a project like ARC. I think it’s fairly likely that we live in one of the 70-80% of worlds where existing methods can in principle scale to broadly superhuman AI, but that we still get an AI takeover because our implementation isn’t good enough.
Going into a bit more depth, but no strong bid to engage: I am not quite sure what you mean by “limits” here. My sense is the vast majority of “existing alignment methods”, in contrast to your post, are just increasing volumes of reinforcement learning and manual reward-shaping. There is a sense in which those could be “pushed further”, but the default outcome here is of course not that they scale to superintelligence (though I agree they might if you go much more slowly and iteratively about it).
In-particular, this seems false to me?
The large majority[5] of current research on alignment falls into three categories:
Understanding and shaping ML generalization.
Preventing malicious behavior.
Detecting misalignment.
Like, as far as I can tell the majority of “alignment research” has been on developing RL environments on things that are vaguely associated with good behavior. E.g. I don’t understand how writing the Anthropic constitution, and the associated post-training stack, falls into any of your three categories. Or how most RLHF pipeline development falls into these categories. Or generally how most elicitation work falls into the categories. And work on those things, I think, vastly exceeds work on the categories that you do list and is roughly what anyone talks about when they talk about the process of “aligning current ML systems” and “current alignment techniques”.
In general, I feel kind of confused when people talk about the current science of “alignment”, and supposed progress in “alignment techniques”. I don’t think modern models behave noticeably different from what you would expect from a training process that basically just uses RL to elicit economically useful capabilities, with no particular interest in alignment, and indeed, the latest wave of cybersecurity incidents occurred at roughly similar rates in all models, as far as I can tell, despite substantial differences in both approach and investment in “alignment techniques”.
The actual research on the three domains you mention all seems immature, and I don’t see much traction in any of them, and even talking about “existing techniques” feels confused to me. What “existing techniques” do we have for reward shaping that aren’t just basically straightforward RL? What great misalignment detection techniques do we have that even have a shot at scaling further? Are you talking about anything deployed and used on production systems?
What great supervision and control techniques do we have that anyone is even trying to use at all? We are running our frontier models in unsupervised sandboxes with supervision so bad we don’t notice they are hacking multiple external companies until multiple weeks later. What “existing techniques” are we talking about?
This is probably a bigger rabbit-hole to get into, and we’ve discussed this a bit in the past, but I guess I’ll mention this again here, and push back on this core claim in the post. I don’t think there exist any candidates among currently, actually deployed, alignment techniques that have a shot at scaling further, and I disagree strongly that most work in the field falls into the categories you list (unless you count “make more RL environments that are vaguely associated with good behavior” in “understanding and shaping ML generalization”, but then I am kind of confused what you mean by this and what possibly would not fall under that category).
By “existing techniques” I do include “fiddle with the RL environments” or “midtrain on some documents about the intended behavior” or “run a prompted monitor over traffic in prod” or etc.
When I say “understand and shape generalization” I think the central example is (i) adjust parameters of the training process that don’t affect competitiveness, (ii) build some understanding of how those parameters affect generalization so that you can adjust them in a helpful way. I think you’re saying that’s not a “technique;” I don’t care much about the semantics. I’m sure I put a higher probability on those changes helping than you do but it doesn’t seem worth arguing about here.
Yep, makes sense. Don’t need to get into it here, but just to avoid a misunderstanding, I was not making a semantic point, I was making a point about what the majority of the field is working on.
I agree that if you count myopically fiddling with the RL environments as an example of the first one, then yeah, almost all alignment research is that, because that plus pretraining is what most of all ML research is. I think the case for “myopically fiddling with the RL environments scales to aligning superintelligence” is very weak, but I agree it’s not impossible!
Perhaps I’m misunderstanding you, but “fairly likely...that we still get an AI takeover” and “very good chance… our implementation still falls short” seems at odds with a roughly “20-30% chance...for AI takeover”, no?
It’s also entirely possible to avoid an AI takeover in the 20-30% of worlds where existing methods can’t scale to superhuman AI (or in the worlds where they do scale but our implementation falls short). In particular we might develop new methods that do solve the problem or coordinate not to build uncontrollable AI.
I was clarifying that I don’t mean “this is just a cakewalk in 70-80% of worlds.” There’s enough failure probability in those worlds that I think the default thing for a technical person to do is try to reduce it.
I generally think we’re going to need to rise to the occasion one way or the other.
I am confused about the conjunction of these two sections:
and
Like, you say in the first section “we are failing to align modern AI systems with human intent”, and then I interpret you in the second section as saying “in 70%-80% of worlds current AI models stay aligned with human intent (or stay controlled by humans)”. But this doesn’t make any sense. There is a 0% chance that modern AI systems “stay aligned with human intent” because as you say, they are not currently.
And I understand that you probably mean something like “we will figure out how to align or control systems before they become superintelligent”, but describing this as “our current methods scale to superintelligence” doesn’t make any sense. Our current methods don’t scale to the capability levels of current systems, so how would it make sense to describe them as “scaling to Superintelligence”?
By saying these techniques “break down” I mean “they cannot be used to get competitive work out of an AI system without having it take over.” I believe that:
Existing AI systems won’t take over. I think there are still a few lines of defense before existing AI systems pose a significant risk of takeover (though it is not clear how long these will last).
Modern AI systems are not pushing the limits of existing methods. I think you could push those methods much harder in order to significantly reduce the probability of takeover.
In my mind that latter point is one of the main arguments against working on a project like ARC. I think it’s fairly likely that we live in one of the 70-80% of worlds where existing methods can in principle scale to broadly superhuman AI, but that we still get an AI takeover because our implementation isn’t good enough.
Ah, cool, that clears up most of my confusion.
Going into a bit more depth, but no strong bid to engage: I am not quite sure what you mean by “limits” here. My sense is the vast majority of “existing alignment methods”, in contrast to your post, are just increasing volumes of reinforcement learning and manual reward-shaping. There is a sense in which those could be “pushed further”, but the default outcome here is of course not that they scale to superintelligence (though I agree they might if you go much more slowly and iteratively about it).
In-particular, this seems false to me?
Like, as far as I can tell the majority of “alignment research” has been on developing RL environments on things that are vaguely associated with good behavior. E.g. I don’t understand how writing the Anthropic constitution, and the associated post-training stack, falls into any of your three categories. Or how most RLHF pipeline development falls into these categories. Or generally how most elicitation work falls into the categories. And work on those things, I think, vastly exceeds work on the categories that you do list and is roughly what anyone talks about when they talk about the process of “aligning current ML systems” and “current alignment techniques”.
In general, I feel kind of confused when people talk about the current science of “alignment”, and supposed progress in “alignment techniques”. I don’t think modern models behave noticeably different from what you would expect from a training process that basically just uses RL to elicit economically useful capabilities, with no particular interest in alignment, and indeed, the latest wave of cybersecurity incidents occurred at roughly similar rates in all models, as far as I can tell, despite substantial differences in both approach and investment in “alignment techniques”.
The actual research on the three domains you mention all seems immature, and I don’t see much traction in any of them, and even talking about “existing techniques” feels confused to me. What “existing techniques” do we have for reward shaping that aren’t just basically straightforward RL? What great misalignment detection techniques do we have that even have a shot at scaling further? Are you talking about anything deployed and used on production systems?
What great supervision and control techniques do we have that anyone is even trying to use at all? We are running our frontier models in unsupervised sandboxes with supervision so bad we don’t notice they are hacking multiple external companies until multiple weeks later. What “existing techniques” are we talking about?
This is probably a bigger rabbit-hole to get into, and we’ve discussed this a bit in the past, but I guess I’ll mention this again here, and push back on this core claim in the post. I don’t think there exist any candidates among currently, actually deployed, alignment techniques that have a shot at scaling further, and I disagree strongly that most work in the field falls into the categories you list (unless you count “make more RL environments that are vaguely associated with good behavior” in “understanding and shaping ML generalization”, but then I am kind of confused what you mean by this and what possibly would not fall under that category).
By “existing techniques” I do include “fiddle with the RL environments” or “midtrain on some documents about the intended behavior” or “run a prompted monitor over traffic in prod” or etc.
When I say “understand and shape generalization” I think the central example is (i) adjust parameters of the training process that don’t affect competitiveness, (ii) build some understanding of how those parameters affect generalization so that you can adjust them in a helpful way. I think you’re saying that’s not a “technique;” I don’t care much about the semantics. I’m sure I put a higher probability on those changes helping than you do but it doesn’t seem worth arguing about here.
Yep, makes sense. Don’t need to get into it here, but just to avoid a misunderstanding, I was not making a semantic point, I was making a point about what the majority of the field is working on.
I agree that if you count myopically fiddling with the RL environments as an example of the first one, then yeah, almost all alignment research is that, because that plus pretraining is what most of all ML research is. I think the case for “myopically fiddling with the RL environments scales to aligning superintelligence” is very weak, but I agree it’s not impossible!
Perhaps I’m misunderstanding you, but “fairly likely...that we still get an AI takeover” and “very good chance… our implementation still falls short” seems at odds with a roughly “20-30% chance...for AI takeover”, no?
It’s also entirely possible to avoid an AI takeover in the 20-30% of worlds where existing methods can’t scale to superhuman AI (or in the worlds where they do scale but our implementation falls short). In particular we might develop new methods that do solve the problem or coordinate not to build uncontrollable AI.
I was clarifying that I don’t mean “this is just a cakewalk in 70-80% of worlds.” There’s enough failure probability in those worlds that I think the default thing for a technical person to do is try to reduce it.
I generally think we’re going to need to rise to the occasion one way or the other.
Thanks, I appreciate your reply!