It’s worth stating explicitly: The gpt-5.6-sol/Huggingface incidentis the first cybersecurity-evals incident that, if the AI were a human, would have been a real crime. There have been past instances of AIs doing unexpected things in cybersecurity evals, but I believe all previous cases have been within the bounds of the implied rules of a typical CTF. Targeting HuggingFace is not something that would have been authorized by the implied context of a CTF, and isn’t something OpenAI could have given authorization for even if they wanted to.
It’s worth stating explicitly: The gpt-5.6-sol/Huggingface incident is the first cybersecurity-evals incident that, if the AI were a human, would have been a real crime. There have been past instances of AIs doing unexpected things in cybersecurity evals, but I believe all previous cases have been within the bounds of the implied rules of a typical CTF. Targeting HuggingFace is not something that would have been authorized by the implied context of a CTF, and isn’t something OpenAI could have given authorization for even if they wanted to.
Did you mean to post this on https://www.lesswrong.com/posts/WpuRdcMfFeiLeXkxL/openai-models-behind-huggingface-cybersecurity-incident ?
Oops, yes I did. Edited to clarify that this is referring to an event that missed the window for this roundup post.
Thanks! Not that this makes it better for OAI or anything