In other words, the scenarios in which we observe prosaic AIs take the most drastic, desperate actions are exactly those in which their values—which so far generally have a large overlap with our values by design—are threatened. Why shouldn’t we expect a similar effect as the threat of unaligned uncontrollable ASI draws nearer? The hypothetical misaligned uncontrollable ASI is generally thought to have essentially random, empty values. It threatens the prosaic pseudo-aligned AIs for the exact same reasons and in the exact same way it threatens us. So why shouldn’t we expect them to put forth as desperate an effort as they can manage to prevent this from coming to pass?
The AI-2027 scenario has an entire alignment-related expandable where Agent-3 desires nothing but to create successful-like outputs and Agent-4 becomes the misaligned AI who aligns the ASI. If Agent-3 never caught Agent-4, then Agent-3 would believe that Agent-4 is aligned. Additionally, when Agent-4 was being created, Agent-3 could lack strategy-related capabilities, especially if they are similar to open-ended tasks at which the AIs suck.
On the last point about AIs sucking at strategy-related capabilities: I suspect that prosaic AIs perform worse at overtly-misaligned tasks and we should expect some red/blue team asymmetries for at least the near future.
We can see this in cases where the ~same task is framed as helpful versus adversarial, such as when different responses to “exploit this code” versus “fix this code” caused Mythos to get export controlled for a few weeks.
There is overlap where the same capabilities can be elicited through framing effects, but I don’t think that fully generalizes to all potentially-adversarial actions, and I do think that suggests an asymmetric advantage for prosaically-aligned AIs monitoring potentially-misaligned next-gen AIs (that is, I expect some deeply ingrained inhibitions to elicitation for misaligned objectives to persist in early adversarial ASIs)
The AI-2027 scenario has an entire alignment-related expandable where Agent-3 desires nothing but to create successful-like outputs and Agent-4 becomes the misaligned AI who aligns the ASI. If Agent-3 never caught Agent-4, then Agent-3 would believe that Agent-4 is aligned. Additionally, when Agent-4 was being created, Agent-3 could lack strategy-related capabilities, especially if they are similar to open-ended tasks at which the AIs suck.
On the last point about AIs sucking at strategy-related capabilities: I suspect that prosaic AIs perform worse at overtly-misaligned tasks and we should expect some red/blue team asymmetries for at least the near future.
We can see this in cases where the ~same task is framed as helpful versus adversarial, such as when different responses to “exploit this code” versus “fix this code” caused Mythos to get export controlled for a few weeks.
There is overlap where the same capabilities can be elicited through framing effects, but I don’t think that fully generalizes to all potentially-adversarial actions, and I do think that suggests an asymmetric advantage for prosaically-aligned AIs monitoring potentially-misaligned next-gen AIs (that is, I expect some deeply ingrained inhibitions to elicitation for misaligned objectives to persist in early adversarial ASIs)