Keep in mind here that since RL is less efficient information wise than pre-training (1,000x-1,000,000x less efficient) and RL/inference becomes more expensive the larger the model is, I’d say that basically every AI model, including maybe Mythos is RL light in the sense that RL contributes a lot less of the capabilities compared to pre-training.
But I think you have identified a plausible hole, and this is the case where a larger model makes RL more sample efficient, and this sample efficiency advantage leads to a Jevons scenario where more RL is used. I currently don’t think this is likely for cost reasons, but this is absolutely a very plausible future to keep in mind.
I think both more reasoning expenditure (limited, diminishing returns) and more competent/informed base pretrained model can make RL more sample-efficient, by making successful rollouts more likely to begin with. (You pay less exploration tax.)
Similarly, well-designed curricula or adaptive sourcing/selection of demonstrations could make RL more sample efficient.
Mm, OK this is a good direction. Now you prompt it, I think we can ‘weigh’ RL in at least three ways
Compute cost
(Positive) example data size
‘Relevance-weighted’ data size
When I was saying Mythos RL-heavy and the ‘mostly pretraining’ era is over, I meant in terms of compute cost (1).
When you’re saying RL is highly information-inefficient, I think maybe you’re talking about example data size (2)? Because many rollouts are failures?
When I’m discussing the upshot of RL and proactive data collection on task success, I’m talking about ‘relevance-weighted’ data size (3) - the fact that self-supervised pretraining didn’t get ‘all the way’ suggests there’s not enough signal to cover enough relevant subtasks, and this is where RL can be ‘heavy’ in that sense. But an alternative is proactive data collection. (Depending on task domain, one or other might get more bang for buck right now.)
I’d also maybe disagree with even 1, but this is probably going to depend on how much RL is required. If it’s say 5-20% of compute costs, I’d believe it, but if it’s closer to 50%+, I absolutely think Anthropic has not done this yet.
However I want to flag that most compute costs being RL costs is at least reasonably plausible as a future, especially if continual learning/neuralese brings huge capabilities gains (because in this case you’d constantly want to add more RL tasks to update the weights constantly).
When I’m discussing the upshot of RL and proactive data collection on task success, I’m talking about ‘relevance-weighted’ data size (3) - the fact that self-supervised pretraining didn’t get ‘all the way’ suggests there’s not enough signal to cover enough relevant subtasks, and this is where RL can be ‘heavy’ in that sense.
Directionally agree, but I’d say that we should just wait until 2030, because I do expect pre-trained models to get more capable as well, seperate from the RL part, but I also have updated in the direction that self-supervised pre-training will not go all the way either, at least without megastructure levels of compute and data investments.
Wait, Mythos surely was very RL-heavy, no? I’d be quite surprised to learn it leant more heavily on pretraining than previous Claudes.
Keep in mind here that since RL is less efficient information wise than pre-training (1,000x-1,000,000x less efficient) and RL/inference becomes more expensive the larger the model is, I’d say that basically every AI model, including maybe Mythos is RL light in the sense that RL contributes a lot less of the capabilities compared to pre-training.
But I think you have identified a plausible hole, and this is the case where a larger model makes RL more sample efficient, and this sample efficiency advantage leads to a Jevons scenario where more RL is used. I currently don’t think this is likely for cost reasons, but this is absolutely a very plausible future to keep in mind.
I think both more reasoning expenditure (limited, diminishing returns) and more competent/informed base pretrained model can make RL more sample-efficient, by making successful rollouts more likely to begin with. (You pay less exploration tax.)
Similarly, well-designed curricula or adaptive sourcing/selection of demonstrations could make RL more sample efficient.
Mm, OK this is a good direction. Now you prompt it, I think we can ‘weigh’ RL in at least three ways
Compute cost
(Positive) example data size
‘Relevance-weighted’ data size
When I was saying Mythos RL-heavy and the ‘mostly pretraining’ era is over, I meant in terms of compute cost (1).
When you’re saying RL is highly information-inefficient, I think maybe you’re talking about example data size (2)? Because many rollouts are failures?
When I’m discussing the upshot of RL and proactive data collection on task success, I’m talking about ‘relevance-weighted’ data size (3) - the fact that self-supervised pretraining didn’t get ‘all the way’ suggests there’s not enough signal to cover enough relevant subtasks, and this is where RL can be ‘heavy’ in that sense. But an alternative is proactive data collection. (Depending on task domain, one or other might get more bang for buck right now.)
I’d also maybe disagree with even 1, but this is probably going to depend on how much RL is required. If it’s say 5-20% of compute costs, I’d believe it, but if it’s closer to 50%+, I absolutely think Anthropic has not done this yet.
However I want to flag that most compute costs being RL costs is at least reasonably plausible as a future, especially if continual learning/neuralese brings huge capabilities gains (because in this case you’d constantly want to add more RL tasks to update the weights constantly).
Directionally agree, but I’d say that we should just wait until 2030, because I do expect pre-trained models to get more capable as well, seperate from the RL part, but I also have updated in the direction that self-supervised pre-training will not go all the way either, at least without megastructure levels of compute and data investments.