We did not explicitly train Mythos Preview to have these capabilities. Rather, they emerged as a downstream consequence of general improvements in code, reasoning, and autonomy. The same improvements that make the model substantially more effective at patching vulnerabilities also make it substantially more effective at exploiting them.
When AI labs say “we did not explicitly train this model to have these capabilities”, I think it should be read as “this jump in capability is not caused by us training for this capability more than we usually do, but as a result of us applying our standard training toolkit at larger scale”. You see a lot of it with OpenAI, where they sometimes say that a model that achieved some impressive math results was “a general model, not a model specialized for math”. But, like, obviously this doesn’t mean that they did not heavily RLVR it on math tasks, right?
I’m not sure, but my guess is that it’s the same here. Anthropic’s standard training setup probably involves a bunch of hacking-like tasks. What the quote says is that Mythos is differentially good at hacking not because Anthropic spent differentially more resources on training it for hacking.
Maybe I misheard something? If they aren’t training directly on exploitation then your plan would make sense, yeah. But I’m very surprised—Irregular, for example, has an Anthropic logo on their website: https://www.irregular.com/about
I mean here they say they didn’t train mythos for example. https://www.anthropic.com/research/mythos-preview
When AI labs say “we did not explicitly train this model to have these capabilities”, I think it should be read as “this jump in capability is not caused by us training for this capability more than we usually do, but as a result of us applying our standard training toolkit at larger scale”. You see a lot of it with OpenAI, where they sometimes say that a model that achieved some impressive math results was “a general model, not a model specialized for math”. But, like, obviously this doesn’t mean that they did not heavily RLVR it on math tasks, right?
I’m not sure, but my guess is that it’s the same here. Anthropic’s standard training setup probably involves a bunch of hacking-like tasks. What the quote says is that Mythos is differentially good at hacking not because Anthropic spent differentially more resources on training it for hacking.
Maybe I misheard something? If they aren’t training directly on exploitation then your plan would make sense, yeah. But I’m very surprised—Irregular, for example, has an Anthropic logo on their website: https://www.irregular.com/about