Were there studies on RL on tasks with p(success)~25% as opposed to p(success)~1.5%? Suppose that the ECI of models trained on the former tasks goes up by 1E-5 per environment, the ECI of models trained on the latter tasks goes up by 1E-4 per environment, while the “misalignment index” increases, respectively, by 1E-6 and 1E-4 per environment. Then a lab willing to increase the amount of compute and environments spent on RL tenfold would produce models where the “misalignment index” is increased by 0.1 vs. 1. Unfortunately, we don’t know how to rule this conjecture in or out by deeper studies, like a wholesale combination of midtraining and RL on not-so-hard tasks...
And what will produce sufficient capability to make up for this?
Were there studies on RL on tasks with p(success)~25% as opposed to p(success)~1.5%? Suppose that the ECI of models trained on the former tasks goes up by 1E-5 per environment, the ECI of models trained on the latter tasks goes up by 1E-4 per environment, while the “misalignment index” increases, respectively, by 1E-6 and 1E-4 per environment. Then a lab willing to increase the amount of compute and environments spent on RL tenfold would produce models where the “misalignment index” is increased by 0.1 vs. 1. Unfortunately, we don’t know how to rule this conjecture in or out by deeper studies, like a wholesale combination of midtraining and RL on not-so-hard tasks...