Right, so you think the weightings won’t move sufficiently toward interactive/RL curricula for another few years at least? I agree that would preserve some incentive (not quite a constraint) toward keeping externalised CoT in token space rather than embeddings, at which point there’s at least a reasonable chance of that remaining broadly legible.
Yes, absent an intelligence explosion by 2028/2030 ala AI 2027/2040 assumes (and here I’m not going to debate how likely that happens here)
The other point I want to make though is that neuralese recurrence and memory is I think inevitable in the longer run, and by 2030-2032 at the latest, incentives start pushing ever more towards neuralese recurrence and memory, so I think plans that rely on CoT legibility are still basically doomed in the medium to long-term.
I’m not sure compute vs data is the way I’d frame this; rather I’d say that the ‘fossil fuel’ of crudely curated scraped and compiled data still has some runway(which weakly constrains toward unsupervised pretraining and externalised CoT architectures as we’ve discussed). If data pipelines start getting more synthetic/RL-heavy, creation/processing/learning of that is still compute-constrained.
Fair point, especially since sample-efficient ML models are likely to require way more compute, at least at inference, and quite plausibly training compute too than current methods, so yeah compute is likely still the bottleneck after a short transition from ‘fossil fuel’ pre-training data to RL/synthetic-data heavy (and emphasis intentionally placed on the RL part), so data being the primary bottleneck is only true for a short time.
Which incidentally answers your question here about why the AI Futures model takes the key input to be compute/effective compute, because under the assumption of RL dominating post-training eventually, the creation of effective datasets for various tasks becomes bottlenecked on compute, rather than data itself.
I honestly don’t know what to forecast in terms of reliance on one or other source of data. There’s probably still a lot of room for dedicated curation of non-interactive curricula, or a kind of partially-interactive (but not live RL) approach like DAgger. But interactive environments and RL are easier to scale and turn the crank on once set up.
My current take is that the labs are in fact doing dedicated curation/partially interactive but not live RL approaches for at least some of RLVR, but I’m of the opinion that outside of a few fields where you can trivially parallelize RL environments like easily-grindable math or programming/SWE jobs, this will largely not work because feedback loops are long and the data is essentially non-stationary/always changing, so you cannot forgo online RL.
So it does work in some case, but it doesn’t really work nearly as well as the labs need, and this is why I expect them to go for live RL approaches.
Also as you say, interactive environments/RL are easier to scale and turn the crank on, so incentives favor online RL even if it isn’t necessary.
Yes, absent an intelligence explosion by 2028/2030 ala AI 2027/2040 assumes (and here I’m not going to debate how likely that happens here)
The other point I want to make though is that neuralese recurrence and memory is I think inevitable in the longer run, and by 2030-2032 at the latest, incentives start pushing ever more towards neuralese recurrence and memory, so I think plans that rely on CoT legibility are still basically doomed in the medium to long-term.
Fair point, especially since sample-efficient ML models are likely to require way more compute, at least at inference, and quite plausibly training compute too than current methods, so yeah compute is likely still the bottleneck after a short transition from ‘fossil fuel’ pre-training data to RL/synthetic-data heavy (and emphasis intentionally placed on the RL part), so data being the primary bottleneck is only true for a short time.
Which incidentally answers your question here about why the AI Futures model takes the key input to be compute/effective compute, because under the assumption of RL dominating post-training eventually, the creation of effective datasets for various tasks becomes bottlenecked on compute, rather than data itself.
My current take is that the labs are in fact doing dedicated curation/partially interactive but not live RL approaches for at least some of RLVR, but I’m of the opinion that outside of a few fields where you can trivially parallelize RL environments like easily-grindable math or programming/SWE jobs, this will largely not work because feedback loops are long and the data is essentially non-stationary/always changing, so you cannot forgo online RL.
So it does work in some case, but it doesn’t really work nearly as well as the labs need, and this is why I expect them to go for live RL approaches.
Also as you say, interactive environments/RL are easier to scale and turn the crank on, so incentives favor online RL even if it isn’t necessary.