Even if multiple actors believe that brain-like architectures are easy to build, I don’t like the strategy of “hope that alignment is easy and race ASAP towards a unipolar world with a surveillance panopticon where people are strictly prohibited from developing unsanctioned ASI”. Having multiple actors actively trying to take over the world and impose surveillance panopticons would be bad! This would cause chaos, shorten timelines, and make even ‘easy’ alignment strategies very difficult to implement. It would also lead to terrible concentrations of power in the event that you win (i.e. take over the world).
In such a world, it would be better to try to collect evidence of whether alignment really is easy, and push governments for very restrictive hardware limitations as outlined here.
One of the many ways we got lucky in the HuggingFace incident is that the agents didn’t encrypt their messages on the message board and delete the keys at the end of the session[1]. If they had done this, it would have been much harder to understand what is going on, though we could have made some progress by looking at the CoTs and action logs. If the agents were careful (e.g. by doing a better job of tampering with the logs), this would have been difficult to understand as well.
I don’t think future misaligned agents will make this mistake, especially if they know there are additional monitors that they need to evade.
Though they did come close when they cryptographically signed their messages