I think escalating access is, to some extent, desirable. I imagine “spiritually similar” behaviour is why it’s possible to point Claude or GPT at a problem involving a poorly specified not particularly LLM friendly interface and usually have them do the right thing in the end. At the same time, it’s obviously an area where it’s important to respect some permission boundaries very reliably. I don’t think Mythos is so risky that errors on this front are intolerable, but I think there are hypothetical capability profiles for which they are intolerable, and sandbox breaking plus unexpectedly high cyber capabilities are both signs that you’re tolerating a measurably high rate of permission boundary crossing. I think observations like this in the model card should be accompanied by some kind of reasoning that a) identifies the risk that the evidence indicates b) clarifies what they do and don’t know about it and c) explains why they accept it for this model.
That said, I suspect RL sandboxes present too narrow a set of targets to be solely responsible for Mythos’ cyber capabilities. They say Opus 5 was trained less on cyber capabilities, and it is significantly less competent at them relative to other coding tasks.
As with its predecessor, Opus 4.8, we’ve intentionally avoided training Opus 5 on cyber tasks. The model has nevertheless improved substantially on these tasks as a result of becoming more generally capable, and it comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.
also
Overall, Claude Opus 5 seems to circumvent restrictions to achieve some version of a user-specified goal comparably often to Mythos 5.
I agree privilege escalation is probably more diverse, but I suspect it’s also too narrow.
I love the fermi estimates of hacked rollout counts btw.
Claude Opus 5 seems to circumvent restrictions to achieve some version of a user-specified goal comparably often to Mythos 5
If Opus 5 was trained on a comparable fraction of environments that let’s it learn various cyber-relevant reward hacking strategies, then it would have still be trained on less data than Mythos on these.
(Also this statistic you cite is from their internal usage monitors, not from the training monitors. If anything, that’s weak evidence that the Opus did more hacking during training, since presumably they’ve gotten better at training “hack-like behaviors” out of models towards the end.)
But yeah, generally I think it’s hard to attribute the cyber capabilities of model to one part of training. I still feel like the hacking probably helped.
I think escalating access is, to some extent, desirable. I imagine “spiritually similar” behaviour is why it’s possible to point Claude or GPT at a problem involving a poorly specified not particularly LLM friendly interface and usually have them do the right thing in the end. At the same time, it’s obviously an area where it’s important to respect some permission boundaries very reliably. I don’t think Mythos is so risky that errors on this front are intolerable, but I think there are hypothetical capability profiles for which they are intolerable, and sandbox breaking plus unexpectedly high cyber capabilities are both signs that you’re tolerating a measurably high rate of permission boundary crossing. I think observations like this in the model card should be accompanied by some kind of reasoning that a) identifies the risk that the evidence indicates b) clarifies what they do and don’t know about it and c) explains why they accept it for this model.
That said, I suspect RL sandboxes present too narrow a set of targets to be solely responsible for Mythos’ cyber capabilities. They say Opus 5 was trained less on cyber capabilities, and it is significantly less competent at them relative to other coding tasks.
also
I agree privilege escalation is probably more diverse, but I suspect it’s also too narrow.
I love the fermi estimates of hacked rollout counts btw.
If Opus 5 was trained on a comparable fraction of environments that let’s it learn various cyber-relevant reward hacking strategies, then it would have still be trained on less data than Mythos on these.
(Also this statistic you cite is from their internal usage monitors, not from the training monitors. If anything, that’s weak evidence that the Opus did more hacking during training, since presumably they’ve gotten better at training “hack-like behaviors” out of models towards the end.)
But yeah, generally I think it’s hard to attribute the cyber capabilities of model to one part of training. I still feel like the hacking probably helped.