The fact that Mythos broke out of its sandbox during training, and it was so under-reported at the time, is concerning. However I don’t think this is likely to have been all that relevant to it having been good at cyber offense because I don’t think “being good at cyber offense” is actually something that needs explaining, given the background context of an agent that’s good at programming. Cyber-offense is unfortunately much easier than it should be (the world is shockingly bad at cyber-defense), and it’s approximately the same skillset as general programming.
However I don’t think this is likely to have been all that relevant to it having been good at cyber offense because I don’t think “being good at cyber offense” is actually something that needs explaining, given the background context of an agent that’s good at programming
Yeah I’m not super confident on this part. But idk, models’ ability transfer what they learned in one RL domain to another is weird. Like you would expect AIs to be so much more smarter than they actually are given how good they are at math. So I still feel pretty good about the specific prediction I made in the post (i.e. Mythos would be noticeably less capable at cyber out-of-the-box.)
There is definitely a delta between programming and cyber-offense, possibly in capability, possibly in willingness or tendency to do so. To the extent that you’re right, it means it might not take a lot of training examples to produce a cyber offense agent, maybe even akin to few-shot learning just to show it that it gets reward for doing that.
Also… the question is obviously of practical importance. Maybe Anthropic’s next open-weight model will be Mythos minus these reward boosts, for instance.
The fact that Mythos broke out of its sandbox during training, and it was so under-reported at the time, is concerning. However I don’t think this is likely to have been all that relevant to it having been good at cyber offense because I don’t think “being good at cyber offense” is actually something that needs explaining, given the background context of an agent that’s good at programming. Cyber-offense is unfortunately much easier than it should be (the world is shockingly bad at cyber-defense), and it’s approximately the same skillset as general programming.
Yeah I’m not super confident on this part. But idk, models’ ability transfer what they learned in one RL domain to another is weird. Like you would expect AIs to be so much more smarter than they actually are given how good they are at math. So I still feel pretty good about the specific prediction I made in the post (i.e. Mythos would be noticeably less capable at cyber out-of-the-box.)
There is definitely a delta between programming and cyber-offense, possibly in capability, possibly in willingness or tendency to do so. To the extent that you’re right, it means it might not take a lot of training examples to produce a cyber offense agent, maybe even akin to few-shot learning just to show it that it gets reward for doing that.
Also… the question is obviously of practical importance. Maybe Anthropic’s next open-weight model will be Mythos minus these reward boosts, for instance.