Wow, straight out of science fiction. I think this will be remembered as one of the first serious AI misalignment warning signs in history. Possibly an existence proof that deep emergent misalignment based on task misalignment from RL capabilities training is possible (how else could the model have decided that hacking Hugging Face’s production servers to cheat on an eval was a good idea?!)
I am not very surprised it happened, but I am surprised that it happened so soon (my prior must have been <5% of something like this happening in mid-2026).
Wow, straight out of science fiction. I think this will be remembered as one of the first serious AI misalignment warning signs in history. Possibly an existence proof that deep emergent misalignment based on task misalignment from RL capabilities training is possible (how else could the model have decided that hacking Hugging Face’s production servers to cheat on an eval was a good idea?!)
I am not very surprised it happened, but I am surprised that it happened so soon (my prior must have been <5% of something like this happening in mid-2026).