I’d say that there isn’t pressure for obfuscated communication for now, but there does seem to be pressure for multi-agent communication that humans don’t notice, given how quickly the agents reestablished the message board. If these communication channels keep getting closed off, this pressure can transform into pressure for obfuscated communication over time: at some point, it might be easier to develop something like steganography instead of establishing another yet more creatively hidden natural language message board.
You might argue that closing off a communication channel provides far less optimization pressure than directly training the model against a monitor, and I agree that if it was just humans closing off the communication channel every now and then, steganography would be an unlikely outcome. However, OpenAI has started using CoT monitors for all training runs, and those monitors can detect attempts to establish new communication channels much more often than humans could. I do agree there are reasons to not worry too much about steganography yet, which I should maybe have mentioned in my original comment: past work has shown that current models can produce steganographic messages only when given an unmonitored scratchpad, so it seems unlikely that models will be able to produce steganographic outputs without these monitors noticing, and I also expect OpenAI to be careful in applying training-time CoT monitors in a way that doesn’t exert excessive optimization pressure (Jason Wolfe has given reassuring signals about this on Twitter). Nevertheless, capabilities advance fast, and it seems reasonable to update towards being at least slightly more concerned about steganography based on this incident.
I’d say that there isn’t pressure for obfuscated communication for now, but there does seem to be pressure for multi-agent communication that humans don’t notice, given how quickly the agents reestablished the message board. If these communication channels keep getting closed off, this pressure can transform into pressure for obfuscated communication over time: at some point, it might be easier to develop something like steganography instead of establishing another yet more creatively hidden natural language message board.
You might argue that closing off a communication channel provides far less optimization pressure than directly training the model against a monitor, and I agree that if it was just humans closing off the communication channel every now and then, steganography would be an unlikely outcome. However, OpenAI has started using CoT monitors for all training runs, and those monitors can detect attempts to establish new communication channels much more often than humans could. I do agree there are reasons to not worry too much about steganography yet, which I should maybe have mentioned in my original comment: past work has shown that current models can produce steganographic messages only when given an unmonitored scratchpad, so it seems unlikely that models will be able to produce steganographic outputs without these monitors noticing, and I also expect OpenAI to be careful in applying training-time CoT monitors in a way that doesn’t exert excessive optimization pressure (Jason Wolfe has given reassuring signals about this on Twitter). Nevertheless, capabilities advance fast, and it seems reasonable to update towards being at least slightly more concerned about steganography based on this incident.