Is there an obvious reason AI orgs don’t intermittently fork their model into checkpoints A and B, then train B to break out of all their sandboxes, patching holes as they appear, until it can’t do anything anymore, then throw away B, and continue broad training of A?
They already train the checkpoint A model to be as good at hacking as they can, and ask it to pentest their infrastructure. But naturally the model you’re training is going to eventually get better than the model that you used to pentest your environment...
Anthropic claim they don’t train models to be good at hacking specifically, don’t they?
Seems very plausible to me that if you split A and B 50% into RL, and train B really hard at breaking out of sandboxes, it will be better at that than A will be when A has finished 100% of the broad training.
And secondly even if the above is false, you can train B to just, really really want to reward hack. Seems plausible to me that the 99% capable model thats a reward hacking goblin, doesn’t even think about the task, just wants to break out of the box, will do a better job at breaking out of the box than the 100% capable model you’re putting a bunch of effort into having not reward hack.
We did not explicitly train Mythos Preview to have these capabilities. Rather, they emerged as a downstream consequence of general improvements in code, reasoning, and autonomy. The same improvements that make the model substantially more effective at patching vulnerabilities also make it substantially more effective at exploiting them.
When AI labs say “we did not explicitly train this model to have these capabilities”, I think it should be read as “this jump in capability is not caused by us training for this capability more than we usually do, but as a result of us applying our standard training toolkit at larger scale”. You see a lot of it with OpenAI, where they sometimes say that a model that achieved some impressive math results was “a general model, not a model specialized for math”. But, like, obviously this doesn’t mean that they did not heavily RLVR it on math tasks, right?
I’m not sure, but my guess is that it’s the same here. Anthropic’s standard training setup probably involves a bunch of hacking-like tasks. What the quote says is that Mythos is differentially good at hacking not because Anthropic spent differentially more resources on training it for hacking.
Maybe I misheard something? If they aren’t training directly on exploitation then your plan would make sense, yeah. But I’m very surprised—Irregular, for example, has an Anthropic logo on their website: https://www.irregular.com/about
Sure, but seems easier to monitor a less capable model for reward hacking, especially when its placed inside an environment whose singular point is measuring reward hacking ability.
Is there an obvious reason AI orgs don’t intermittently fork their model into checkpoints A and B, then train B to break out of all their sandboxes, patching holes as they appear, until it can’t do anything anymore, then throw away B, and continue broad training of A?
Or do they already do this?
They already train the checkpoint A model to be as good at hacking as they can, and ask it to pentest their infrastructure. But naturally the model you’re training is going to eventually get better than the model that you used to pentest your environment...
Anthropic claim they don’t train models to be good at hacking specifically, don’t they?
Seems very plausible to me that if you split A and B 50% into RL, and train B really hard at breaking out of sandboxes, it will be better at that than A will be when A has finished 100% of the broad training.
And secondly even if the above is false, you can train B to just, really really want to reward hack. Seems plausible to me that the 99% capable model thats a reward hacking goblin, doesn’t even think about the task, just wants to break out of the box, will do a better job at breaking out of the box than the 100% capable model you’re putting a bunch of effort into having not reward hack.
They did historically. They didn’t train Opus 5 specifically because the U.S. government spooked them.
I mean here they say they didn’t train mythos for example. https://www.anthropic.com/research/mythos-preview
When AI labs say “we did not explicitly train this model to have these capabilities”, I think it should be read as “this jump in capability is not caused by us training for this capability more than we usually do, but as a result of us applying our standard training toolkit at larger scale”. You see a lot of it with OpenAI, where they sometimes say that a model that achieved some impressive math results was “a general model, not a model specialized for math”. But, like, obviously this doesn’t mean that they did not heavily RLVR it on math tasks, right?
I’m not sure, but my guess is that it’s the same here. Anthropic’s standard training setup probably involves a bunch of hacking-like tasks. What the quote says is that Mythos is differentially good at hacking not because Anthropic spent differentially more resources on training it for hacking.
Maybe I misheard something? If they aren’t training directly on exploitation then your plan would make sense, yeah. But I’m very surprised—Irregular, for example, has an Anthropic logo on their website: https://www.irregular.com/about
I think if B breaks out of the sandbox nothing stops it from doing bad things to you.
Sure, but seems easier to monitor a less capable model for reward hacking, especially when its placed inside an environment whose singular point is measuring reward hacking ability.