I think it would be safest to just restart the RL run with the patched environments. If you train against only the instances of hacking that your monitor can detect, you risk indirectly training the model to evade the monitor.
They’re almost certainly not doing it for Mythos? This result is from reviewing what had already happening in the training run. It’ll be way too expensive to patch everything and retrain the model.
I think it would be safest to just restart the RL run with the patched environments. If you train against only the instances of hacking that your monitor can detect, you risk indirectly training the model to evade the monitor.
Anthropic is the safety-first frontier AI lab, surely they’re doing at least that much.
They’re almost certainly not doing it for Mythos? This result is from reviewing what had already happening in the training run. It’ll be way too expensive to patch everything and retrain the model.
Sorry that was sarcasm. Poe’s law, I know.
Sure, but how much would that cost in time and money? I think it’s important to recognize the choices devs are facing.