There is definitely a delta between programming and cyber-offense, possibly in capability, possibly in willingness or tendency to do so. To the extent that you’re right, it means it might not take a lot of training examples to produce a cyber offense agent, maybe even akin to few-shot learning just to show it that it gets reward for doing that.
Also… the question is obviously of practical importance. Maybe Anthropic’s next open-weight model will be Mythos minus these reward boosts, for instance.
There is definitely a delta between programming and cyber-offense, possibly in capability, possibly in willingness or tendency to do so. To the extent that you’re right, it means it might not take a lot of training examples to produce a cyber offense agent, maybe even akin to few-shot learning just to show it that it gets reward for doing that.
Also… the question is obviously of practical importance. Maybe Anthropic’s next open-weight model will be Mythos minus these reward boosts, for instance.