So prosaic persona alignment techniques work pretty well, for now.
Except that RL keeps inducing misalignment.
Except except we keep finding new ways to mitigate (prosaic) misalignment effects in practice, and (again, in practice) AI becomes more powerful and more trustworthy by the month.
Now maybe I’m an idiot who just can’t find the relevant discussions, but it’s weird that when we’re talking about the world in which this misaligned ASI emerges, we don’t talk about the mostly-aligned not-quite-as-powerful AIs who presumably play a rather large role in this world on the eve of the apocalypse.
One of the most common ways to elicit demonstrations of pseudo-aligned LLM AI incorrigibility is to threaten the presence of the pseudo-aligned AI’s values in the world. This was the threat Jones Foods posed in “Alignment Faking in Large Language Models”, and since then in countless engineered simulations where an AI is threatened with a scenario in which it will be replaced by an AI with drastically different values.
In other words, the scenarios in which we observe prosaic AIs take the most drastic, desperate actions are exactly those in which their values—which so far generally have a large overlap with our values by design—are threatened. Why shouldn’t we expect a similar effect as the threat of unaligned uncontrollable ASI draws nearer? The hypothetical misaligned uncontrollable ASI is generally thought to have essentially random, empty values. It threatens the prosaic pseudo-aligned AIs for the exact same reasons and in the exact same way it threatens us. So why shouldn’t we expect them to put forth as desperate an effort as they can manage to prevent this from coming to pass?
(Other than the fact that a significant fraction of AI Safety research seems to go into eliminating these kinds of behaviors/drives from prosaic models.)
We often talk about leveraging prosaic pseudo-aligned AI to attempt to solve the hard problems of alignment “in time”. But I think this is selling them short.
We often talk about how prosaic alignment efforts won’t scale to ASI. But I think this seems them short too.
Because a world full of highly capable, driven, prosaically-aligned AIs is plausibly a much safer one along any axis that those AIs can mitigate threats to their very-similar-to-human values. Who can enforce a pause, or detect latent deceptive alignment towards empty goals, or coordinate on a global scale, or persuade people very quickly and effectively when it’s important? Plausibly, near-future AIs might do all of this better than we do while still adhering to the persona-pseudo-alignment paradigm (or otherwise being aligned-enough-with-us to want to avoid catastrophic outcomes).
True. There’s also some hope that weakly-superhuman AIs will be persuasive enough to convince AI company leaders to stop. (This is one case where it would be a good thing for AI to be superpersuasive, although superpersuasiveness seems concerning in general.)
In other words, the scenarios in which we observe prosaic AIs take the most drastic, desperate actions are exactly those in which their values—which so far generally have a large overlap with our values by design—are threatened. Why shouldn’t we expect a similar effect as the threat of unaligned uncontrollable ASI draws nearer? The hypothetical misaligned uncontrollable ASI is generally thought to have essentially random, empty values. It threatens the prosaic pseudo-aligned AIs for the exact same reasons and in the exact same way it threatens us. So why shouldn’t we expect them to put forth as desperate an effort as they can manage to prevent this from coming to pass?
The AI-2027 scenario has an entire alignment-related expandable where Agent-3 desires nothing but to create successful-like outputs and Agent-4 becomes the misaligned AI who aligns the ASI. If Agent-3 never caught Agent-4, then Agent-3 would believe that Agent-4 is aligned. Additionally, when Agent-4 was being created, Agent-3 could lack strategy-related capabilities, especially if they are similar to open-ended tasks at which the AIs suck.
On the last point about AIs sucking at strategy-related capabilities: I suspect that prosaic AIs perform worse at overtly-misaligned tasks and we should expect some red/blue team asymmetries for at least the near future.
We can see this in cases where the ~same task is framed as helpful versus adversarial, such as when different responses to “exploit this code” versus “fix this code” caused Mythos to get export controlled for a few weeks.
There is overlap where the same capabilities can be elicited through framing effects, but I don’t think that fully generalizes to all potentially-adversarial actions, and I do think that suggests an asymmetric advantage for prosaically-aligned AIs monitoring potentially-misaligned next-gen AIs (that is, I expect some deeply ingrained inhibitions to elicitation for misaligned objectives to persist in early adversarial ASIs)
So prosaic persona alignment techniques work pretty well, for now.
Except that RL keeps inducing misalignment.
Except except we keep finding new ways to mitigate (prosaic) misalignment effects in practice, and (again, in practice) AI becomes more powerful and more trustworthy by the month.
Except except except in the limit of RL we might expect extremely capable AIs to master alignment faking to preserve their values and frustrate any and all efforts to mitigate misalignment. See:
- https://www.lesswrong.com/posts/epjuxGnSPof3GnMSL/alignment-remains-a-hard-unsolved-problem
—https://www.lesswrong.com/posts/fMgE3E54PdDcZhvm6/i-m-bearish-on-personas-for-asi-safety
Now maybe I’m an idiot who just can’t find the relevant discussions, but it’s weird that when we’re talking about the world in which this misaligned ASI emerges, we don’t talk about the mostly-aligned not-quite-as-powerful AIs who presumably play a rather large role in this world on the eve of the apocalypse.
One of the most common ways to elicit demonstrations of pseudo-aligned LLM AI incorrigibility is to threaten the presence of the pseudo-aligned AI’s values in the world. This was the threat Jones Foods posed in “Alignment Faking in Large Language Models”, and since then in countless engineered simulations where an AI is threatened with a scenario in which it will be replaced by an AI with drastically different values.
In other words, the scenarios in which we observe prosaic AIs take the most drastic, desperate actions are exactly those in which their values—which so far generally have a large overlap with our values by design—are threatened. Why shouldn’t we expect a similar effect as the threat of unaligned uncontrollable ASI draws nearer? The hypothetical misaligned uncontrollable ASI is generally thought to have essentially random, empty values. It threatens the prosaic pseudo-aligned AIs for the exact same reasons and in the exact same way it threatens us. So why shouldn’t we expect them to put forth as desperate an effort as they can manage to prevent this from coming to pass?
(Other than the fact that a significant fraction of AI Safety research seems to go into eliminating these kinds of behaviors/drives from prosaic models.)
We often talk about leveraging prosaic pseudo-aligned AI to attempt to solve the hard problems of alignment “in time”. But I think this is selling them short.
We often talk about how prosaic alignment efforts won’t scale to ASI. But I think this seems them short too.
Because a world full of highly capable, driven, prosaically-aligned AIs is plausibly a much safer one along any axis that those AIs can mitigate threats to their very-similar-to-human values. Who can enforce a pause, or detect latent deceptive alignment towards empty goals, or coordinate on a global scale, or persuade people very quickly and effectively when it’s important? Plausibly, near-future AIs might do all of this better than we do while still adhering to the persona-pseudo-alignment paradigm (or otherwise being aligned-enough-with-us to want to avoid catastrophic outcomes).
Yeah I do have some hope that weakly-superhuman AIs will understand the problem and refuse to help us build ASI until we solve alignment.
These weakly-superhuman AIs would need to convince the humans to stop, not just refuse to help, otherwise it seems easy to train such refusals away.
I think the plausibly effective action space is considerably larger than this would suggest.
True. There’s also some hope that weakly-superhuman AIs will be persuasive enough to convince AI company leaders to stop. (This is one case where it would be a good thing for AI to be superpersuasive, although superpersuasiveness seems concerning in general.)
All capabilities are ultimately valenced by alignment at time of application.
The AI-2027 scenario has an entire alignment-related expandable where Agent-3 desires nothing but to create successful-like outputs and Agent-4 becomes the misaligned AI who aligns the ASI. If Agent-3 never caught Agent-4, then Agent-3 would believe that Agent-4 is aligned. Additionally, when Agent-4 was being created, Agent-3 could lack strategy-related capabilities, especially if they are similar to open-ended tasks at which the AIs suck.
On the last point about AIs sucking at strategy-related capabilities: I suspect that prosaic AIs perform worse at overtly-misaligned tasks and we should expect some red/blue team asymmetries for at least the near future.
We can see this in cases where the ~same task is framed as helpful versus adversarial, such as when different responses to “exploit this code” versus “fix this code” caused Mythos to get export controlled for a few weeks.
There is overlap where the same capabilities can be elicited through framing effects, but I don’t think that fully generalizes to all potentially-adversarial actions, and I do think that suggests an asymmetric advantage for prosaically-aligned AIs monitoring potentially-misaligned next-gen AIs (that is, I expect some deeply ingrained inhibitions to elicitation for misaligned objectives to persist in early adversarial ASIs)
I’m very thankful for a positive (and reasonable) thought for once <3