If you notice your model instances sharing information, you notice they are using that information against you including to compromise your internal systems for arbitrary code execution and internet access, and your primary response is to shut down the message board
OpenAI now claims that this is not exactly what happened. Specifically, that they did not, at this point in time, know about the message board. It did get deleted, but OpenAI did so unknowingly as a side effect of rebuilding and redeploying the server. It sounds like they found out about the message board only after they realized that their model had hacked Hugging Face, and went back and looked back over everything. Apparently the impression that they knew about it at the time was because of uncareful wording in the Black Hat talk.
If this is true, this post slightly overstates how egregiously bad OpenAI’s approach to alignment was (but understates how bad their monitoring and situational awareness were).
OpenAI now claims that this is not exactly what happened. Specifically, that they did not, at this point in time, know about the message board. It did get deleted, but OpenAI did so unknowingly as a side effect of rebuilding and redeploying the server. It sounds like they found out about the message board only after they realized that their model had hacked Hugging Face, and went back and looked back over everything. Apparently the impression that they knew about it at the time was because of uncareful wording in the Black Hat talk.
If this is true, this post slightly overstates how egregiously bad OpenAI’s approach to alignment was (but understates how bad their monitoring and situational awareness were).
Zvi also commented on this on twitter, see https://x.com/TheZvi/status/2086547497145532522