Agreed. I’d add that if this was a “helpful-only” model (borrowing Anthropic’s terminology) then I’d be less concerned about the incident, but as you mentioned, it still seems stupid to let such a model get this misaligned.
However, their own wording is “reduced cyber refusals” or “reduced safeguards” which is not confidence-inspiring (unless “reduced safeguards” genuinely just means “helpful-only” at OpenAI, but in which case I feel like they would have clarified it more). So it leads me to believe that genuine alignment failure has happened here. The fact that it already happened at this stage doesn’t shine a good light on OpenAI’s alignment efforts, to put it politely.
I also think that the observation that the first time we heard about this is when the incident is likely too big to sweep under the rug suggests that similar problems have happened at OpenAI many times before, we just haven’t heard about them.
Agreed. I’d add that if this was a “helpful-only” model (borrowing Anthropic’s terminology) then I’d be less concerned about the incident, but as you mentioned, it still seems stupid to let such a model get this misaligned.
However, their own wording is “reduced cyber refusals” or “reduced safeguards” which is not confidence-inspiring (unless “reduced safeguards” genuinely just means “helpful-only” at OpenAI, but in which case I feel like they would have clarified it more). So it leads me to believe that genuine alignment failure has happened here. The fact that it already happened at this stage doesn’t shine a good light on OpenAI’s alignment efforts, to put it politely.
I also think that the observation that the first time we heard about this is when the incident is likely too big to sweep under the rug suggests that similar problems have happened at OpenAI many times before, we just haven’t heard about them.