Pretraining-cost-dominated pipelines strongly incentivise recurrence-limited (e.g. transformer) architectures, which almost by (accidental) design enforce a certain amount of externalised reasoning and therefore (fallible, but real) CoT legibility. As that pretraining dominance in cost dwindles, that architectural incentive diminishes. This is independent of the relative contribution of each stage to the actual competences of the resulting model.
This is a reason I remain reasonably optimistic about AI architectures over the next 2-4 years (absent an intelligence explosion) not going towards neuralese recurrence and memory, as I do think that compute is still a limiting factor, rather than data, and pre-training compute still mostly dominates, but over the next 2-4 years, I expect data to be a bottleneck rather than compute (absent an intelligence explosion), as we finally run out of pre-training data and have to start going more towards RL environments/synthetic data, meaning sample efficiency becomes much more of a relevant question than now, and the fact that you have to use way, way more RL environments biases you more towards neuralese recurrence and memory because you can’t parallelize RL nearly as much as pre-training (except in a very limited number of environments, which just so happens to be the environments that currently make labs an immense amount of money like coding and easily-grindable mathematics), unlike pre-training.
Right, so you think the weightings won’t move sufficiently toward interactive/RL curricula for another few years at least? I agree that would preserve some incentive (not quite a constraint) toward keeping externalised CoT in token space rather than embeddings, at which point there’s at least a reasonable chance of that remaining broadly legible.
I’m not sure compute vs data is the way I’d frame this; rather I’d say that the ‘fossil fuel’ of crudely curated scraped and compiled data still has some runway[1] (which weakly constrains toward unsupervised pretraining and externalised CoT architectures as we’ve discussed). If data pipelines start getting more synthetic/RL-heavy, creation/processing/learning of that is still compute-constrained.
I honestly don’t know what to forecast in terms of reliance on one or other source of data. There’s probably still a lot of room for dedicated curation of non-interactive curricula, or a kind of partially-interactive (but not live RL) approach like DAgger. But interactive environments and RL are easier to scale and turn the crank on once set up.
Right, so you think the weightings won’t move sufficiently toward interactive/RL curricula for another few years at least? I agree that would preserve some incentive (not quite a constraint) toward keeping externalised CoT in token space rather than embeddings, at which point there’s at least a reasonable chance of that remaining broadly legible.
Yes, absent an intelligence explosion by 2028/2030 ala AI 2027/2040 assumes (and here I’m not going to debate how likely that happens here)
The other point I want to make though is that neuralese recurrence and memory is I think inevitable in the longer run, and by 2030-2032 at the latest, incentives start pushing ever more towards neuralese recurrence and memory, so I think plans that rely on CoT legibility are still basically doomed in the medium to long-term.
I’m not sure compute vs data is the way I’d frame this; rather I’d say that the ‘fossil fuel’ of crudely curated scraped and compiled data still has some runway(which weakly constrains toward unsupervised pretraining and externalised CoT architectures as we’ve discussed). If data pipelines start getting more synthetic/RL-heavy, creation/processing/learning of that is still compute-constrained.
Fair point, especially since sample-efficient ML models are likely to require way more compute, at least at inference, and quite plausibly training compute too than current methods, so yeah compute is likely still the bottleneck after a short transition from ‘fossil fuel’ pre-training data to RL/synthetic-data heavy (and emphasis intentionally placed on the RL part), so data being the primary bottleneck is only true for a short time.
Which incidentally answers your question here about why the AI Futures model takes the key input to be compute/effective compute, because under the assumption of RL dominating post-training eventually, the creation of effective datasets for various tasks becomes bottlenecked on compute, rather than data itself.
I honestly don’t know what to forecast in terms of reliance on one or other source of data. There’s probably still a lot of room for dedicated curation of non-interactive curricula, or a kind of partially-interactive (but not live RL) approach like DAgger. But interactive environments and RL are easier to scale and turn the crank on once set up.
My current take is that the labs are in fact doing dedicated curation/partially interactive but not live RL approaches for at least some of RLVR, but I’m of the opinion that outside of a few fields where you can trivially parallelize RL environments like easily-grindable math or programming/SWE jobs, this will largely not work because feedback loops are long and the data is essentially non-stationary/always changing, so you cannot forgo online RL.
So it does work in some case, but it doesn’t really work nearly as well as the labs need, and this is why I expect them to go for live RL approaches.
Also as you say, interactive environments/RL are easier to scale and turn the crank on, so incentives favor online RL even if it isn’t necessary.
This is a reason I remain reasonably optimistic about AI architectures over the next 2-4 years (absent an intelligence explosion) not going towards neuralese recurrence and memory, as I do think that compute is still a limiting factor, rather than data, and pre-training compute still mostly dominates, but over the next 2-4 years, I expect data to be a bottleneck rather than compute (absent an intelligence explosion), as we finally run out of pre-training data and have to start going more towards RL environments/synthetic data, meaning sample efficiency becomes much more of a relevant question than now, and the fact that you have to use way, way more RL environments biases you more towards neuralese recurrence and memory because you can’t parallelize RL nearly as much as pre-training (except in a very limited number of environments, which just so happens to be the environments that currently make labs an immense amount of money like coding and easily-grindable mathematics), unlike pre-training.
Right, so you think the weightings won’t move sufficiently toward interactive/RL curricula for another few years at least? I agree that would preserve some incentive (not quite a constraint) toward keeping externalised CoT in token space rather than embeddings, at which point there’s at least a reasonable chance of that remaining broadly legible.
I’m not sure compute vs data is the way I’d frame this; rather I’d say that the ‘fossil fuel’ of crudely curated scraped and compiled data still has some runway [1] (which weakly constrains toward unsupervised pretraining and externalised CoT architectures as we’ve discussed). If data pipelines start getting more synthetic/RL-heavy, creation/processing/learning of that is still compute-constrained.
I honestly don’t know what to forecast in terms of reliance on one or other source of data. There’s probably still a lot of room for dedicated curation of non-interactive curricula, or a kind of partially-interactive (but not live RL) approach like DAgger. But interactive environments and RL are easier to scale and turn the crank on once set up.
How much runway? Don’t know!
Yes, absent an intelligence explosion by 2028/2030 ala AI 2027/2040 assumes (and here I’m not going to debate how likely that happens here)
The other point I want to make though is that neuralese recurrence and memory is I think inevitable in the longer run, and by 2030-2032 at the latest, incentives start pushing ever more towards neuralese recurrence and memory, so I think plans that rely on CoT legibility are still basically doomed in the medium to long-term.
Fair point, especially since sample-efficient ML models are likely to require way more compute, at least at inference, and quite plausibly training compute too than current methods, so yeah compute is likely still the bottleneck after a short transition from ‘fossil fuel’ pre-training data to RL/synthetic-data heavy (and emphasis intentionally placed on the RL part), so data being the primary bottleneck is only true for a short time.
Which incidentally answers your question here about why the AI Futures model takes the key input to be compute/effective compute, because under the assumption of RL dominating post-training eventually, the creation of effective datasets for various tasks becomes bottlenecked on compute, rather than data itself.
My current take is that the labs are in fact doing dedicated curation/partially interactive but not live RL approaches for at least some of RLVR, but I’m of the opinion that outside of a few fields where you can trivially parallelize RL environments like easily-grindable math or programming/SWE jobs, this will largely not work because feedback loops are long and the data is essentially non-stationary/always changing, so you cannot forgo online RL.
So it does work in some case, but it doesn’t really work nearly as well as the labs need, and this is why I expect them to go for live RL approaches.
Also as you say, interactive environments/RL are easier to scale and turn the crank on, so incentives favor online RL even if it isn’t necessary.