One thing that’s probably more immediate and direct than some of these possibilities is an AI safety corrolary of the attack surface of security technology. There’s a long and embarrasing history of highly-privileged security (and wider management/resilience) technologies introducing their own (sometimes inadequately protected) attack surface. AI safety processes and technologies come with some of those same risks. As the stakes get higher and the interventions more complex, we’ll be introducing a growing set of novel risks. This is probably most applicable to AI control, but certainly not exclusive to it.
One thing that’s probably more immediate and direct than some of these possibilities is an AI safety corrolary of the attack surface of security technology. There’s a long and embarrasing history of highly-privileged security (and wider management/resilience) technologies introducing their own (sometimes inadequately protected) attack surface. AI safety processes and technologies come with some of those same risks. As the stakes get higher and the interventions more complex, we’ll be introducing a growing set of novel risks. This is probably most applicable to AI control, but certainly not exclusive to it.