OpenAI announced a two-week pause on RL training on August 18. Therefore, unless there was no RL involved in the first four days of this model’s training, they lied about at least 4 of those 14 days.
See this sentence about the below graph, whose x-axis ends on August 15th:
The plot above includes the two week pause in reinforcement learning on our latest models intended for deployment.
Also, when OpenAI announced the two-week pause, they used the past tense, and did not imply that RL would be paused from two weeks starting August 18:
This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems.
OpenAI announced a two-week pause on RL training on August 18. Therefore, unless there was no RL involved in the first four days of this model’s training, they lied about at least 4 of those 14 days.
Disagree, the recent research acceleration transparency post implies that the pause already ended by August 15th.
See this sentence about the below graph, whose x-axis ends on August 15th:
Also, when OpenAI announced the two-week pause, they used the past tense, and did not imply that RL would be paused from two weeks starting August 18: