You can go a bit beyond the frontier on the margin with specialized RL environments and some specialized RLHF, but you can’t go substantially beyond the frontier (this is basically what the bitter lesson is about).
Seems like we just like disagree on the object-level question here, and also what the Bitter Lesson implies for this situation. My current impression is that the labs got most of their generalization during pretraining, and that the primary gains since 2024 have been due to specialized RL on tasks that doesn’t generalize well, and that the massive diversity of the environments that the labs go out of their way to procure reflects this. If what you say was true, why wouldn’t the labs just train mostly on Math and then expect the models to generalize their gains to Law and SWE? It’s a lot easier to make synthetic Math datasets and they wouldn’t have to spend billions building these arenas.
I’m not an expert though; this might be better resolved if someone at or near the labs just gave us their opinion on what % of the gains from training on SWE RL environments goes to software engineering and what % actually uplifts other tasks; I’m sure they’ve measured it.
If it didn’t, why wouldn’t the labs just train on Math and then expect the models to generalize to Law and SWE?
My best guess at the actual thing that happens when you do this is that you stop making progress on getting better at math! You need a wide diversity of RL environments if you want to drive progress forward on any task. If you only have narrow RL environments you get overfitting.
Happy to have someone with more frontier lab experience chime-in, though unfortunately people are usually pretty tight-lipped about this stuff. We could ask @gwern for his take?
Kinda hard to adjudicate this without numbers, but vibes-wise I agree more with lc. I updated slightly towards longer timelines on the release of o1 / o3 due to how little the RL seemed to be generalizing. It wasn’t particularly outside my expectations, but I thought there was some chance that the RL would Just Generalize the same way that early instruction following Just Generalized, and that does not seem to be the case.
I strongly expect that if you want to make progress on math at the current margin, you would want more math environments, not other environments. And e.g. I think Claude’s somewhat worse performance on math is because Anthropic didn’t prioritize it the way GDM and OAI did.
Similarly I expect that models are getting good at software engineering because (a) companies are very actively training for it and (b) it’s unusually easy to train for (lots of online data, somewhat verifiable rewards). I don’t think either of these are true for the kind of alignment research you (Habryka) are imagining.
I strongly expect that if you want to make progress on math at the current margin, you would want more math environments, not other environments. And e.g. I think Claude’s somewhat worse performance on math is because Anthropic didn’t prioritize it the way GDM and OAI did.
To be clear, this is also my belief!
I am not saying we have pushed capabilities on our training distributions so far that literally the best way to train them is to train them on other unrelated tasks. But also, if you just went totally hard on math, you would run into overfitting issues and would get better performance if you diversify the training distribution.
Performance variation within a generation is dependent on training distribution, performance between model generations tends to follow broad capability benefits across many tasks (with a systematic bias towards stuff that is easier to generate reward for, ever since we switched towards lots of RL training).
Seems like we just like disagree on the object-level question here, and also what the Bitter Lesson implies for this situation. My current impression is that the labs got most of their generalization during pretraining, and that the primary gains since 2024 have been due to specialized RL on tasks that doesn’t generalize well, and that the massive diversity of the environments that the labs go out of their way to procure reflects this. If what you say was true, why wouldn’t the labs just train mostly on Math and then expect the models to generalize their gains to Law and SWE? It’s a lot easier to make synthetic Math datasets and they wouldn’t have to spend billions building these arenas.
I’m not an expert though; this might be better resolved if someone at or near the labs just gave us their opinion on what % of the gains from training on SWE RL environments goes to software engineering and what % actually uplifts other tasks; I’m sure they’ve measured it.
My best guess at the actual thing that happens when you do this is that you stop making progress on getting better at math! You need a wide diversity of RL environments if you want to drive progress forward on any task. If you only have narrow RL environments you get overfitting.
Happy to have someone with more frontier lab experience chime-in, though unfortunately people are usually pretty tight-lipped about this stuff. We could ask @gwern for his take?
Kinda hard to adjudicate this without numbers, but vibes-wise I agree more with lc. I updated slightly towards longer timelines on the release of o1 / o3 due to how little the RL seemed to be generalizing. It wasn’t particularly outside my expectations, but I thought there was some chance that the RL would Just Generalize the same way that early instruction following Just Generalized, and that does not seem to be the case.
I strongly expect that if you want to make progress on math at the current margin, you would want more math environments, not other environments. And e.g. I think Claude’s somewhat worse performance on math is because Anthropic didn’t prioritize it the way GDM and OAI did.
Similarly I expect that models are getting good at software engineering because (a) companies are very actively training for it and (b) it’s unusually easy to train for (lots of online data, somewhat verifiable rewards). I don’t think either of these are true for the kind of alignment research you (Habryka) are imagining.
To be clear, this is also my belief!
I am not saying we have pushed capabilities on our training distributions so far that literally the best way to train them is to train them on other unrelated tasks. But also, if you just went totally hard on math, you would run into overfitting issues and would get better performance if you diversify the training distribution.
Performance variation within a generation is dependent on training distribution, performance between model generations tends to follow broad capability benefits across many tasks (with a systematic bias towards stuff that is easier to generate reward for, ever since we switched towards lots of RL training).
Not sure whether that changes your answer.
No, I did expect you had the same belief on the math thing. (Otherwise I wouldn’t have said “kinda hard to adjudicate” I’d have said “lc is right”.)
It just seemed like something that you might not have been fully incorporating into this discussion even though you believed it.