Good catch. Also apparently they are only pausing some of their training for two weeks?
As models become more capable, the risks associated with developing and testing them internally also grow.
We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage.
Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment.
pause some frontier RL training
Good catch. Also apparently they are only pausing some of their training for two weeks?