Pretraining scaling should be calibrated using the observed differences between models of different sizes, trained with different amounts of compute. The end of the current trend in rapid scaling of pretraining is dictated by running out of compute or pretraining data, and there’s plausibly 200T tokens of unique data for a 2e29 FLOPs model of 2031 with 30x sparsity, which is effectively just 4x undertrained (uses a 4x lower D/N ratio than would be compute optimal), so it’s not much different from a compute optimally trained 2e29 FLOPs model. At that point, finding 10x more compute will be complicated, it’s not practical to improve on 30x sparsity very far, and in any case that requires bigger scale-up systems that will take a few more years.
This is probably a significantly bigger difference than between Sonnet 5 and Mythos 5 (or between GPT-5.4 and Astra), so a model that’s this far above Mythos 5 (or Astra) seems sufficient (with some redundancy) as a base model for RLVRing automated model training skills. That in turn enables automated training of all other skills that the contemporary revision of the model sufficiently comprehends to write RLVR tasks/environments/graders for, and this process of prosaic RSI is what I expect to pass the AGI milestone in the sense of unbounded eventual progress. But since only RLVR is instilling novel deep skills in this process, and it’s not so far above Mythos 5 (and probably Astra) that it starts making impossible leaps of insight, it’s not necessarily even faster than humans at conceptual research (even though it’s very likely capable of it over sufficiently long time, across sufficiently many iterations of training the next model).
Pretraining scaling should be calibrated using the observed differences between models of different sizes, trained with different amounts of compute. The end of the current trend in rapid scaling of pretraining is dictated by running out of compute or pretraining data, and there’s plausibly 200T tokens of unique data for a 2e29 FLOPs model of 2031 with 30x sparsity, which is effectively just 4x undertrained (uses a 4x lower D/N ratio than would be compute optimal), so it’s not much different from a compute optimally trained 2e29 FLOPs model. At that point, finding 10x more compute will be complicated, it’s not practical to improve on 30x sparsity very far, and in any case that requires bigger scale-up systems that will take a few more years.
To calibrate expectations about the 2031 model, Mythos 5 is plausibly a 1.3e27 FLOPs model with 8x sparsity, while Opus 4.5+ is plausibly a 3e26 FLOPs model with 4x sparsity. Every 2x of sparsity increases effective compute about 1.4x. Thus the difference between Opus 4.5+ and Mythos 5 is about 6x in effective compute, and the difference between Mythos 5 and the 2031 model is 300x in effective compute, 3.3x as far on the logarithmic scale (ignoring the slight undertraining effect, and the worse training data quality when more data is needed).
This is probably a significantly bigger difference than between Sonnet 5 and Mythos 5 (or between GPT-5.4 and Astra), so a model that’s this far above Mythos 5 (or Astra) seems sufficient (with some redundancy) as a base model for RLVRing automated model training skills. That in turn enables automated training of all other skills that the contemporary revision of the model sufficiently comprehends to write RLVR tasks/environments/graders for, and this process of prosaic RSI is what I expect to pass the AGI milestone in the sense of unbounded eventual progress. But since only RLVR is instilling novel deep skills in this process, and it’s not so far above Mythos 5 (and probably Astra) that it starts making impossible leaps of insight, it’s not necessarily even faster than humans at conceptual research (even though it’s very likely capable of it over sufficiently long time, across sufficiently many iterations of training the next model).