Is the rise of RL because we’ve run out of non-synthetic data with which to scale pretraining?
I’d guess that the rise of RL is because this is the ~best way (in terms of compute efficiency + scalability) to keep pushing frontier capabilities, especially in domains that seem high value
Is the rise of RL because we’ve run out of non-synthetic data with which to scale pretraining?
I’d guess that the rise of RL is because this is the ~best way (in terms of compute efficiency + scalability) to keep pushing frontier capabilities, especially in domains that seem high value