So I agree that “hacking Anthropic” usually implies gaining access to parts of Anthropic’s networks/systems that the model is not supposed to have, and Mythos preview (probably) did not achieve that.
However, the model did gain unauthorized access somewhere. It repeatedly broke into some part of its sandboxed environment. An environment created by Anthropic (or its suppliers). An environment deliberately designed to contain its actions. An environment that the model nonetheless “hacked.” I think it’s reasonable to say that the model’s hacking behaviors are directed at Anthropic.
So saying something like “Mythos preview acquired greater permissions within its sandboxed RL environment” almost makes you forget that humans built these sandboxes and did not want them broken. While Mythos Preview’s hacking did not cause Anthropic any direct damage, it definitely reduced the model’s usefulness and harmed Anthropic indirectly. In general, I am worried that there’s a tendency for lab employees to sanitize the language they use around misaligned model behaviors and make things feel less crazy and insane than they are.
But I guess allowing people to go “a-ha! you made a mistake. It didn’t really hack Anthropic!” distracts from the point I’m trying to make. So I’ve changed the title of the post to say “Anthropic’s sandboxes” instead of Anthropic.
So I agree that “hacking Anthropic” usually implies gaining access to parts of Anthropic’s networks/systems that the model is not supposed to have, and Mythos preview (probably) did not achieve that.
However, the model did gain unauthorized access somewhere. It repeatedly broke into some part of its sandboxed environment. An environment created by Anthropic (or its suppliers). An environment deliberately designed to contain its actions. An environment that the model nonetheless “hacked.” I think it’s reasonable to say that the model’s hacking behaviors are directed at Anthropic.
So saying something like “Mythos preview acquired greater permissions within its sandboxed RL environment” almost makes you forget that humans built these sandboxes and did not want them broken. While Mythos Preview’s hacking did not cause Anthropic any direct damage, it definitely reduced the model’s usefulness and harmed Anthropic indirectly. In general, I am worried that there’s a tendency for lab employees to sanitize the language they use around misaligned model behaviors and make things feel less crazy and insane than they are.
But I guess allowing people to go “a-ha! you made a mistake. It didn’t really hack Anthropic!” distracts from the point I’m trying to make. So I’ve changed the title of the post to say “Anthropic’s sandboxes” instead of Anthropic.