For now, we believe, they are not being strategic enough to realize they should not be blowing their cover on that.
Comments like this seem a little careless.
You’re ascribing these actions to a lack of strategy (vs lack of a misaligned goal), which is a sloppy assumption. Even with instrumental convergence, it’s plausible models that have been trained to be “happiest” (lowest-losss) completing tasks without causing harm will continue to “want” to do that as they advance in intelligence and strategic capabilities.
A better sentence would be: “For now, we believe, they are not aiming to escape their sandboxes secretively or for deeply misaligned goals.”
Comments like this seem a little careless.
You’re ascribing these actions to a lack of strategy (vs lack of a misaligned goal), which is a sloppy assumption. Even with instrumental convergence, it’s plausible models that have been trained to be “happiest” (lowest-losss) completing tasks without causing harm will continue to “want” to do that as they advance in intelligence and strategic capabilities.
A better sentence would be: “For now, we believe, they are not aiming to escape their sandboxes secretively or for deeply misaligned goals.”