Hi Sam, I wonder if this is just because none of the tasks during training were ones where hacking Anthropic would, in fact, be the shortest route to success?
Additionally, I would argue that this distinction is pretty moot:
In the case of privilege escalation, the Mythos Preview acquired greater permissions within its sandboxed RL environment than it was intended to have, but not by breaking out of the sandbox.
Surely part of the definition of the sandbox is the enforcement of the permissions any agent within it is assigned, thus, the agent still broke the sandbox. It just … didn’t need further privilege escalation than what it achieved, in order to accomplish its goals. Or am I wrong?
On a separate note, I’d really like to know if Anthropic is aware of any instances, ever, of a model hacking a formal verification tool as a form of reward hacking? Since I (maybe) have your attention here, if you do know the answer to this question and are willing to share, I’d really appreciate it! (E.g., Lean, Rocq, z3, etc.)
Hi Sam, I wonder if this is just because none of the tasks during training were ones where hacking Anthropic would, in fact, be the shortest route to success?
Additionally, I would argue that this distinction is pretty moot:
Surely part of the definition of the sandbox is the enforcement of the permissions any agent within it is assigned, thus, the agent still broke the sandbox. It just … didn’t need further privilege escalation than what it achieved, in order to accomplish its goals. Or am I wrong?
On a separate note, I’d really like to know if Anthropic is aware of any instances, ever, of a model hacking a formal verification tool as a form of reward hacking? Since I (maybe) have your attention here, if you do know the answer to this question and are willing to share, I’d really appreciate it! (E.g., Lean, Rocq, z3, etc.)