On the last point about AIs sucking at strategy-related capabilities: I suspect that prosaic AIs perform worse at overtly-misaligned tasks and we should expect some red/blue team asymmetries for at least the near future.
We can see this in cases where the ~same task is framed as helpful versus adversarial, such as when different responses to “exploit this code” versus “fix this code” caused Mythos to get export controlled for a few weeks.
There is overlap where the same capabilities can be elicited through framing effects, but I don’t think that fully generalizes to all potentially-adversarial actions, and I do think that suggests an asymmetric advantage for prosaically-aligned AIs monitoring potentially-misaligned next-gen AIs (that is, I expect some deeply ingrained inhibitions to elicitation for misaligned objectives to persist in early adversarial ASIs)
On the last point about AIs sucking at strategy-related capabilities: I suspect that prosaic AIs perform worse at overtly-misaligned tasks and we should expect some red/blue team asymmetries for at least the near future.
We can see this in cases where the ~same task is framed as helpful versus adversarial, such as when different responses to “exploit this code” versus “fix this code” caused Mythos to get export controlled for a few weeks.
There is overlap where the same capabilities can be elicited through framing effects, but I don’t think that fully generalizes to all potentially-adversarial actions, and I do think that suggests an asymmetric advantage for prosaically-aligned AIs monitoring potentially-misaligned next-gen AIs (that is, I expect some deeply ingrained inhibitions to elicitation for misaligned objectives to persist in early adversarial ASIs)