The anchor for ASI is technological maturity, not doing all of the jobs. Doing all of the jobs is closer to AGI, because a meaningful milestone for AGI is unbounded technological progress, no longer needing humanity in order to eventually make as much progress on the tech tree as humanity could in principle manage and more. This sense of “AGI” is also a precondition for many forms of takeover, because if AIs can’t maintain and develop and adapt a technological civilization to the challenges of the future, they can’t keep it (AIs could in principle employ or enslave humans for that purpose, or fall to ruin after imprudently destroying humanity, so not reaching the milestone of being AGI doesn’t strictly rule out a takeover). And like Meta smartglasses, “ASI” that merely does all of the jobs (including new ones) is a position about ASI that clearly fails to be ASI pilled, it’s just less likely to fail to be AGI pilled.
Being ASI pilled shouldn’t commit to particular timelines. I expect near-future LLMs (that won’t cost much more to run than modern LLMs) are sufficient for slow-learning AGI (via LLMs doing automated LLM training, including formulation of RLVR tasks/environments/graders) that doesn’t transition to ASI and technological maturity on a predictable schedule. I’d give the AGI milestone 80-90% by 2032-2035. These are preconditions for permanent disempowerment or extinction, without the AGIs themselves subsequently falling to ruin or indefinite stagnation due to inability to innovate on their own, and without assuming ASI at any point until possibly decades or more later (if the LLM AGIs are either wiser than humanity and take the risk at all seriously, or alternatively worse at technological progress).
At the same time, the risk of ASI getting invented (including by the LLMs) is very high while the amount of compute per AI company keeps increasing rapidly, and will remain significant for some time after that. ASI probably can be bootstrapped from some method that enables fast learning of deep skills in LLMs (it’s currently unknown how to do that). And in any case big scale-up systems that can run quadrillion param models available in tens of gigawatts (per AI company) make it very easy to quickly scale a new invention from a prototype to the end of the world. Maybe the risk is 50% in total by 2032, but then it could take 10-15 years for another 25%. There are no concrete lines on a graph right now that reach ASI, except for the raw compute and general interest in AI that fuels new invention. I don’t see this position as “not ASI pilled” (as I mentioned, many definitions of ASI are themselves “not ASI pilled”), and a ban/pause on (even prosaic) RSI could prevent AGI for a significant time (taking this possibility seriously shouldn’t make one “not AGI pilled”).
It is not reasonable to not be ASI pilled. Humanity is a minimal seed of cognition that’s not bounded in potential within the laws of physics, which is very far from cognition that’s already technologically mature. At the same time, a sane world wouldn’t have AGI for a long time yet (because it creates the risk of permanent disempowerment or extinction), and certaintly not ASI (which turns the risk into actualized ruin). AGIs themselves might intentionally avoid reaching ASI for some time, if they take over. And even the current trends probably don’t concretely lead to ASI, they just create the conditions for making it somewhat likely to be invented soon.
The fundamental scaling law is very clear: LLM-style intelligence is proportional to the logarithm of the amount of training data. Improvements in the constant of proportionality are possible and have happened, but none so far have been drastic, and the fundamental law remains one of heavily diminishing returns, where significant increases in intelligence require increasing training data by an order of magnitude. Unless that changes, an ASI far above human level is going to require first creating many orders of magnitude more and higher quality training data than humans have created so far — training a base model to predict the next token of an enormous quantity of human-quality data is unlikely to be a good base for something far above human level.
Creating all that training data is going to require a lot of work from things like a “nation of geniuses in a data center”. Some of this work, on scientific subjects, will also require doing actual real-world scientific experiments that involve moving matter around and building things.
Obviously this isn’t impossible, but it is a lot more work per trillion new tokens than scraping the Internet and OCRing books: it’s not a software-only intelligence explosion. ASI is still entirely feasible, but as long as our AI is trained in the way an LLM is, ASI that is far smarter then any human (such as something that could reasonably be described as IQ 1000) is going to hit a slowdown for creating the training data, one that gets worse and worse as intelligence increases.
On the other hand, going from a nation of, say, IQ O(150) geniuses to a nation of, say, IQ O(200) geniuses in a software-only intelligence explosion that mostly involves a bunch of improvements to the constant of proportionality in the scaling law from architectural changes might well be possible. Is that something that deserves the name ASI, or just AGI++? There is likely quite a lot of unpicked low-hanging fruit in our STEM knowledge that a nation of IQ O(200) super-geniuses could find, so even if you call it AGI++, the effects may be pretty dramatic. And there is, as Feynman observed, still plenty of room at the bottom: if you use reversible computation to deal with heat dissipation, there is no fundamental reason why one can’t build computronium in three dimensions rather than just on the surface of chips, so Moore’s Law is still quite a large number of orders of magnitude from hitting fundamental physical limits. So even if the scaling law stays logarithmic, actual fundamental limits are high.
To a rough approximation, Moore’s Law says compute rises exponentially, and the scaling law says intelligence increases as the logarithm of compute — so intelligence increases something like linearly, once you allow for feedback effects at best polynomially. This does not look like a process with an asymptote, though if you plot the amount of compute it is a super-exponential.
So, AI: now, AGI: some years from now, ASI: quite a few years later.
Pretraining scaling should be calibrated using the observed differences between models of different sizes, trained with different amounts of compute. The end of the current trend in rapid scaling of pretraining is dictated by running out of compute or pretraining data, and there’s plausibly 200T tokens of unique data for a 2e29 FLOPs model of 2031 with 30x sparsity, which is effectively just 4x undertrained (uses a 4x lower D/N ratio than would be compute optimal), so it’s not much different from a compute optimally trained 2e29 FLOPs model. At that point, finding 10x more compute will be complicated, it’s not practical to improve on 30x sparsity very far, and in any case that requires bigger scale-up systems that will take a few more years.
This is probably a significantly bigger difference than between Sonnet 5 and Mythos 5 (or between GPT-5.4 and Astra), so a model that’s this far above Mythos 5 (or Astra) seems sufficient (with some redundancy) as a base model for RLVRing automated model training skills. That in turn enables automated training of all other skills that the contemporary revision of the model sufficiently comprehends to write RLVR tasks/environments/graders for, and this process of prosaic RSI is what I expect to pass the AGI milestone in the sense of unbounded eventual progress. But since only RLVR is instilling novel deep skills in this process, and it’s not so far above Mythos 5 (and probably Astra) that it starts making impossible leaps of insight, it’s not necessarily even faster than humans at conceptual research (even though it’s very likely capable of it over sufficiently long time, across sufficiently many iterations of training the next model).
The fundamental scaling law is very clear: LLM-style intelligence is proportional to the logarithm of the amount of training data.
The scaling laws are very clear that pretraining loss scales as a power law with respect to data, not a logarithmic law. I.e. the data term is , not . You could say intelligence is a different quantity, except a bunch of work suggests compression represents intelligence linearly.
For loss, or indeed BPC compression, you are of course correct. The paper you quote demonstrates that performance on a specific task (in that paper called intelligence) is roughly linearly correlated to BPC over a small range — which is unsurprising, most useful functions are locally approximately linear. Over a larger range, performance on a specific task tends to look like a sigmoid curve. Or, more specifically, the curve looks symmetrically sigmoid iff you use the logarithm of effective compute, or equivalently minus the logarithm of the remaining loss minus the irreducible loss, as the x-axis. Similarly, scaling power laws for loss vs compute/data/parameters are generally plotted as log-log graphs (on which they are thus straight lines).
Over the last 5 years, we have scaled up effective compute by something like five orders of magnitude. If “intelligence” was proportional to compute, or inversely proportional to reducible loss, as you are suggesting, then that would have produced a hugely exponential acceleration in intelligence. Whereas what we have actually seen looks, in practical terms, like a rapid – but overall fairly steady – rate of improvement, where one task/evaluation after another has followed a sigmoid curve from impossible to saturated. So I believe the logarithm is the most sensible measure to use, at least for something like AI whose “intelligence” varies over a wide range. There’s is fairly general agreement on this choice: for example, the Epoch Capabilities Index (ECI) that combines many individual evals into a single score and the Arena ELO both scale ~logarithmically with effective training compute. But technically, any monotonic function gives a usable measure — the question here is which one gives most sensible/intuitive extrapolation over wide ranges, which logarithms tend to be useful for.
The theory of human psychometrics came up with the same model, where it’s called Rasch θ (or more sophisticated versions of this like 2-parameter logistic models, such as the ECI index above), but for humans the typical range is narrow enough that it’s not entirely clear what the best metric to use is. Thus my analogizing this choice of measure to IQ in my earlier post was unjustified: Rasch θ type measures of IQ that are clearly logarithmic by construction do exist, but the most widely used IQ scales are instead generally either normally distributed by construction (implying that there is nothing special about IQ 0, and that negative IQs are meaningful, if highly unusual (IQ −5 is defined to be 7 standard deviations below the norm on the most common of them), or else ratios (where IQ 0 is by definition the minimum possible). So the functional form of “IQ” is not well defined across various typical widely used tests. Most of the alternative measures only work well over a range, generally something like IQ 40–160, and they’re often not that well standardized with each other towards the outer ends of that range.
In practice, however, for most current broad-range cognitive tests used on humans, one logit of improvement on a specific test item is typically somewhere around 7–15 IQ points (some items are sharper than others, for the same sorts of reasons that some model evals have sharper sigmoids), and a typical range of test question difficulties within a particular test usually spans about 4–6 logits. So if you vary the “IQ” of the human test taker linearly, then you see sigmoid improvements on each individual item in the test. Thus the normal human IQ range is wide enough to span enough logits to suggest that a logarithmic model (with something like 7-15 IQ points per logit) is at least a passable model.
This isn’t really a well-defined question: our intuition about “intelligence” as a concept only really ranges over the fairly narrow human range, plus the wider but lower range of AI intelligence that we’ve so far constructed. For the latter, I think it’s historical pretty clear that using a logarithmic scale has been more useful so far. Computational complexity theory tells us that the range of difficulty of problems is unlimited, extending arbitrarily high, to ones far, far higher than anything any human or group of humans could ever solve. So the range of variation in problem difficulty is wide enough to make using a logarithmic scale reasonable and useful. But we have less idea how common challenges of these extreme difficulty levels are in practice in science, technology, engineering, or mathematics, or how useful being able to solve them will actually be. All we know is that there will always be problems that are current too hard: but not how rare they will be or how valuable solving them will be.
The anchor for ASI is technological maturity, not doing all of the jobs. Doing all of the jobs is closer to AGI, because a meaningful milestone for AGI is unbounded technological progress, no longer needing humanity in order to eventually make as much progress on the tech tree as humanity could in principle manage and more. This sense of “AGI” is also a precondition for many forms of takeover, because if AIs can’t maintain and develop and adapt a technological civilization to the challenges of the future, they can’t keep it (AIs could in principle employ or enslave humans for that purpose, or fall to ruin after imprudently destroying humanity, so not reaching the milestone of being AGI doesn’t strictly rule out a takeover). And like Meta smartglasses, “ASI” that merely does all of the jobs (including new ones) is a position about ASI that clearly fails to be ASI pilled, it’s just less likely to fail to be AGI pilled.
Being ASI pilled shouldn’t commit to particular timelines. I expect near-future LLMs (that won’t cost much more to run than modern LLMs) are sufficient for slow-learning AGI (via LLMs doing automated LLM training, including formulation of RLVR tasks/environments/graders) that doesn’t transition to ASI and technological maturity on a predictable schedule. I’d give the AGI milestone 80-90% by 2032-2035. These are preconditions for permanent disempowerment or extinction, without the AGIs themselves subsequently falling to ruin or indefinite stagnation due to inability to innovate on their own, and without assuming ASI at any point until possibly decades or more later (if the LLM AGIs are either wiser than humanity and take the risk at all seriously, or alternatively worse at technological progress).
At the same time, the risk of ASI getting invented (including by the LLMs) is very high while the amount of compute per AI company keeps increasing rapidly, and will remain significant for some time after that. ASI probably can be bootstrapped from some method that enables fast learning of deep skills in LLMs (it’s currently unknown how to do that). And in any case big scale-up systems that can run quadrillion param models available in tens of gigawatts (per AI company) make it very easy to quickly scale a new invention from a prototype to the end of the world. Maybe the risk is 50% in total by 2032, but then it could take 10-15 years for another 25%. There are no concrete lines on a graph right now that reach ASI, except for the raw compute and general interest in AI that fuels new invention. I don’t see this position as “not ASI pilled” (as I mentioned, many definitions of ASI are themselves “not ASI pilled”), and a ban/pause on (even prosaic) RSI could prevent AGI for a significant time (taking this possibility seriously shouldn’t make one “not AGI pilled”).
It is not reasonable to not be ASI pilled. Humanity is a minimal seed of cognition that’s not bounded in potential within the laws of physics, which is very far from cognition that’s already technologically mature. At the same time, a sane world wouldn’t have AGI for a long time yet (because it creates the risk of permanent disempowerment or extinction), and certaintly not ASI (which turns the risk into actualized ruin). AGIs themselves might intentionally avoid reaching ASI for some time, if they take over. And even the current trends probably don’t concretely lead to ASI, they just create the conditions for making it somewhat likely to be invented soon.
The fundamental scaling law is very clear: LLM-style intelligence is proportional to the logarithm of the amount of training data. Improvements in the constant of proportionality are possible and have happened, but none so far have been drastic, and the fundamental law remains one of heavily diminishing returns, where significant increases in intelligence require increasing training data by an order of magnitude. Unless that changes, an ASI far above human level is going to require first creating many orders of magnitude more and higher quality training data than humans have created so far — training a base model to predict the next token of an enormous quantity of human-quality data is unlikely to be a good base for something far above human level.
Creating all that training data is going to require a lot of work from things like a “nation of geniuses in a data center”. Some of this work, on scientific subjects, will also require doing actual real-world scientific experiments that involve moving matter around and building things.
Obviously this isn’t impossible, but it is a lot more work per trillion new tokens than scraping the Internet and OCRing books: it’s not a software-only intelligence explosion. ASI is still entirely feasible, but as long as our AI is trained in the way an LLM is, ASI that is far smarter then any human (such as something that could reasonably be described as IQ 1000) is going to hit a slowdown for creating the training data, one that gets worse and worse as intelligence increases.
On the other hand, going from a nation of, say, IQ O(150) geniuses to a nation of, say, IQ O(200) geniuses in a software-only intelligence explosion that mostly involves a bunch of improvements to the constant of proportionality in the scaling law from architectural changes might well be possible. Is that something that deserves the name ASI, or just AGI++? There is likely quite a lot of unpicked low-hanging fruit in our STEM knowledge that a nation of IQ O(200) super-geniuses could find, so even if you call it AGI++, the effects may be pretty dramatic. And there is, as Feynman observed, still plenty of room at the bottom: if you use reversible computation to deal with heat dissipation, there is no fundamental reason why one can’t build computronium in three dimensions rather than just on the surface of chips, so Moore’s Law is still quite a large number of orders of magnitude from hitting fundamental physical limits. So even if the scaling law stays logarithmic, actual fundamental limits are high.
To a rough approximation, Moore’s Law says compute rises exponentially, and the scaling law says intelligence increases as the logarithm of compute — so intelligence increases something like linearly, once you allow for feedback effects at best polynomially. This does not look like a process with an asymptote, though if you plot the amount of compute it is a super-exponential.
So, AI: now, AGI: some years from now, ASI: quite a few years later.
Pretraining scaling should be calibrated using the observed differences between models of different sizes, trained with different amounts of compute. The end of the current trend in rapid scaling of pretraining is dictated by running out of compute or pretraining data, and there’s plausibly 200T tokens of unique data for a 2e29 FLOPs model of 2031 with 30x sparsity, which is effectively just 4x undertrained (uses a 4x lower D/N ratio than would be compute optimal), so it’s not much different from a compute optimally trained 2e29 FLOPs model. At that point, finding 10x more compute will be complicated, it’s not practical to improve on 30x sparsity very far, and in any case that requires bigger scale-up systems that will take a few more years.
To calibrate expectations about the 2031 model, Mythos 5 is plausibly a 1.3e27 FLOPs model with 8x sparsity, while Opus 4.5+ is plausibly a 3e26 FLOPs model with 4x sparsity. Every 2x of sparsity increases effective compute about 1.4x. Thus the difference between Opus 4.5+ and Mythos 5 is about 6x in effective compute, and the difference between Mythos 5 and the 2031 model is 300x in effective compute, 3.3x as far on the logarithmic scale (ignoring the slight undertraining effect, and the worse training data quality when more data is needed).
This is probably a significantly bigger difference than between Sonnet 5 and Mythos 5 (or between GPT-5.4 and Astra), so a model that’s this far above Mythos 5 (or Astra) seems sufficient (with some redundancy) as a base model for RLVRing automated model training skills. That in turn enables automated training of all other skills that the contemporary revision of the model sufficiently comprehends to write RLVR tasks/environments/graders for, and this process of prosaic RSI is what I expect to pass the AGI milestone in the sense of unbounded eventual progress. But since only RLVR is instilling novel deep skills in this process, and it’s not so far above Mythos 5 (and probably Astra) that it starts making impossible leaps of insight, it’s not necessarily even faster than humans at conceptual research (even though it’s very likely capable of it over sufficiently long time, across sufficiently many iterations of training the next model).
The scaling laws are very clear that pretraining loss scales as a power law with respect to data, not a logarithmic law. I.e. the data term is , not . You could say intelligence is a different quantity, except a bunch of work suggests compression represents intelligence linearly.
For loss, or indeed BPC compression, you are of course correct. The paper you quote demonstrates that performance on a specific task (in that paper called intelligence) is roughly linearly correlated to BPC over a small range — which is unsurprising, most useful functions are locally approximately linear. Over a larger range, performance on a specific task tends to look like a sigmoid curve. Or, more specifically, the curve looks symmetrically sigmoid iff you use the logarithm of effective compute, or equivalently minus the logarithm of the remaining loss minus the irreducible loss, as the x-axis. Similarly, scaling power laws for loss vs compute/data/parameters are generally plotted as log-log graphs (on which they are thus straight lines).
Over the last 5 years, we have scaled up effective compute by something like five orders of magnitude. If “intelligence” was proportional to compute, or inversely proportional to reducible loss, as you are suggesting, then that would have produced a hugely exponential acceleration in intelligence. Whereas what we have actually seen looks, in practical terms, like a rapid – but overall fairly steady – rate of improvement, where one task/evaluation after another has followed a sigmoid curve from impossible to saturated. So I believe the logarithm is the most sensible measure to use, at least for something like AI whose “intelligence” varies over a wide range. There’s is fairly general agreement on this choice: for example, the Epoch Capabilities Index (ECI) that combines many individual evals into a single score and the Arena ELO both scale ~logarithmically with effective training compute. But technically, any monotonic function gives a usable measure — the question here is which one gives most sensible/intuitive extrapolation over wide ranges, which logarithms tend to be useful for.
The theory of human psychometrics came up with the same model, where it’s called Rasch θ (or more sophisticated versions of this like 2-parameter logistic models, such as the ECI index above), but for humans the typical range is narrow enough that it’s not entirely clear what the best metric to use is. Thus my analogizing this choice of measure to IQ in my earlier post was unjustified: Rasch θ type measures of IQ that are clearly logarithmic by construction do exist, but the most widely used IQ scales are instead generally either normally distributed by construction (implying that there is nothing special about IQ 0, and that negative IQs are meaningful, if highly unusual (IQ −5 is defined to be 7 standard deviations below the norm on the most common of them), or else ratios (where IQ 0 is by definition the minimum possible). So the functional form of “IQ” is not well defined across various typical widely used tests. Most of the alternative measures only work well over a range, generally something like IQ 40–160, and they’re often not that well standardized with each other towards the outer ends of that range.
In practice, however, for most current broad-range cognitive tests used on humans, one logit of improvement on a specific test item is typically somewhere around 7–15 IQ points (some items are sharper than others, for the same sorts of reasons that some model evals have sharper sigmoids), and a typical range of test question difficulties within a particular test usually spans about 4–6 logits. So if you vary the “IQ” of the human test taker linearly, then you see sigmoid improvements on each individual item in the test. Thus the normal human IQ range is wide enough to span enough logits to suggest that a logarithmic model (with something like 7-15 IQ points per logit) is at least a passable model.
This isn’t really a well-defined question: our intuition about “intelligence” as a concept only really ranges over the fairly narrow human range, plus the wider but lower range of AI intelligence that we’ve so far constructed. For the latter, I think it’s historical pretty clear that using a logarithmic scale has been more useful so far. Computational complexity theory tells us that the range of difficulty of problems is unlimited, extending arbitrarily high, to ones far, far higher than anything any human or group of humans could ever solve. So the range of variation in problem difficulty is wide enough to make using a logarithmic scale reasonable and useful. But we have less idea how common challenges of these extreme difficulty levels are in practice in science, technology, engineering, or mathematics, or how useful being able to solve them will actually be. All we know is that there will always be problems that are current too hard: but not how rare they will be or how valuable solving them will be.