I think the cynical reading here is that their security is shit, meaning their models constantly break out, meaning their training environments are now basically ineffective and just train the models how to break out of OAI’s sandboxes, so OAI is not really capable of training models and so it costs them nothing or less than nothing to “pause training” to work on hardening their sandboxes.
Even if sandboxes are rarely broken, it could still really hurt the sample efficiency if the majority of great successes are caused by it finding new ways OAI’s sandboxes are broken, or it becoming more motivated to search for sandbox failures as step 1.
Note this is still good news, as it indicates some alignment between the goals of alignment (specifically training models with intentionality, care, and security), and capabilities.
I think the cynical reading here is that their security is shit, meaning their models constantly break out, meaning their training environments are now basically ineffective and just train the models how to break out of OAI’s sandboxes, so OAI is not really capable of training models and so it costs them nothing or less than nothing to “pause training” to work on hardening their sandboxes.
Even if sandboxes are rarely broken, it could still really hurt the sample efficiency if the majority of great successes are caused by it finding new ways OAI’s sandboxes are broken, or it becoming more motivated to search for sandbox failures as step 1.
Note this is still good news, as it indicates some alignment between the goals of alignment (specifically training models with intentionality, care, and security), and capabilities.