accumulation of crystallized intelligence that was produced by their own fluid intelligence
Essentially this isn’t happening right now. I think it very likely starts happening within a few years, but will be slow, maybe slower than humanity, and as a result won’t be able to quickly fix the problem of being slow.
“actual” AGI—the kind that probably doesn’t already exist—the kind that has fluid intelligence and AI advantages for recursive self-improvement, which together make it likely to take over the world shortly after being created
Humanity should count as actual AGI, but it remains slow. Similarly, I think LLMs that do prosaic RSI of automatically building the next model (crucially including formulation of new RL tasks/environments/graders) will be actual AGI in the sense that they can (on their own, without humanity’s input) eventually generate and accumulate as much crystallized intelligence as humanity would.
Yet there is a bottleneck of the speed at which they learn novel deep skills (generate and crystallize new pieces of intelligence), going from new RL tasks to new models to new teaching moments that inspire new tasks for the new models that are ready to use the opportunity. This keeps the process at the slow pace of building models rather than at the fast pace of generating tokens, and so AI advantages don’t quickly snowball into fast RSI that changes the fundamental nature of the AI.
Eventually fast RSI happens, but it plausibly takes many years of slow-learning prosaic RSI (for it) to figure out how to make that happen, and humans could end up being faster at figuring it out. As with any basic research, there is no trend that meaningfully predicts how many years it takes. All this time, there’s actual AGI of slow-learning model building that can figure out anything eventually, but only slowly. This AGI might be the lesser factor of danger in triggering fast RSI, compared to the vast amount of compute its usefulness finances, which enables ambitious experiments and rapid scaling of prototypes.
We have large gaps in performance between AIs and humans (sample complexity for learning, ability to generate novel good concepts)
Pretraining very sample inefficiently reconstructs cognitive skills that left evidence about their nature in the text. RL training can use relatively short formulations of tasks/environments/graders to generate even novel pieces of cognitive skills needed to solve the tasks (that the authors of the tasks didn’t necessarily have, as they didn’t necessarily know how to solve the tasks, or how to do so efficiently). In this way, RL training is both sample efficient (with respect to the short formulations of tasks, and the tiny insights each task promotes), and a way of producing novel pieces of intelligence that can be crystallized in the new models.
What’s currently missing is automated formulation of new RL tasks/environments/graders in response to gaps in existing crystallized intelligence (in the current model) with respect to the situations/problems it comes in contact with. I estimate that LLMs of 2031 will be more capable than Mythos 5 by about 3x as much as Mythos 5 is more capable than Opus 4.5+ (strictly before Opus 5). This is likely a gap notably wider than the Sonnet-Mythos gap, above the current frontier models that are already in a weight class capable of producing strong technical results and spontaneous competentcyberoffense. Taking that further step to the models of 2031 is probably insufficient for leaps of insight that quickly show how to make RSI go fast. But this is very likely enough for a base model of 2031 to be sufficiently teachable (using RL tasks) to end up learning how to formulate new RL tasks/environments/graders on its own, starting the slow-learning prosaic RSI process of building the next model automatically.
Thanks. (I have read this with interest, but don’t have much of an overall reply.) A couple comments:
Humanity should count as actual AGI, but it remains slow.
It’s a weird middle ground, where it is fooming but pretty slowly. (I mean, as a human, I would think that, i.e. I would experience time on a scale that’s somewhat faster than humanity’s foom.) And it kinda only half has the A in AGI. If it fully had the A (or I guess, ISGI, in silico general intelligence) then it would foom much faster.
Similarly, I think LLMs that do prosaic RSI of automatically building the next model (crucially including formulation of new RL tasks/environments/graders) will be actual AGI in the sense that they can (on their own, without humanity’s input) eventually generate and accumulate as much crystallized intelligence as humanity would.
Maybe kinda. But my guess is that this would be kind of like calling [hominid evolution by natural selection on genetic variation] an AGI, or calling [the ecosystem of the Earth throughout all of time] an AGI. I mean it would probably be extremely faster & scarier than evolution, but still.
Pretraining very sample inefficiently reconstructs cognitive skills that left evidence about their nature in the text.
It reconstructs a lot of them to a significant extent, but very much misses a lot of them to a significant extent, I think.
But this is very likely enough for a base model of 2031 to be sufficiently teachable (using RL tasks) to end up learning how to formulate new RL tasks/environments/graders on its own, starting the slow-learning prosaic RSI process of building the next model automatically.
This seems plausible-ish. I would weakly expect a lot of plateaus, e.g. due to highly correlated taste, but not strongly and I haven’t thought about it much. (Maybe you’re pricing that in to “slow”.)
Essentially this isn’t happening right now. I think it very likely starts happening within a few years, but will be slow, maybe slower than humanity, and as a result won’t be able to quickly fix the problem of being slow.
Humanity should count as actual AGI, but it remains slow. Similarly, I think LLMs that do prosaic RSI of automatically building the next model (crucially including formulation of new RL tasks/environments/graders) will be actual AGI in the sense that they can (on their own, without humanity’s input) eventually generate and accumulate as much crystallized intelligence as humanity would.
Yet there is a bottleneck of the speed at which they learn novel deep skills (generate and crystallize new pieces of intelligence), going from new RL tasks to new models to new teaching moments that inspire new tasks for the new models that are ready to use the opportunity. This keeps the process at the slow pace of building models rather than at the fast pace of generating tokens, and so AI advantages don’t quickly snowball into fast RSI that changes the fundamental nature of the AI.
Eventually fast RSI happens, but it plausibly takes many years of slow-learning prosaic RSI (for it) to figure out how to make that happen, and humans could end up being faster at figuring it out. As with any basic research, there is no trend that meaningfully predicts how many years it takes. All this time, there’s actual AGI of slow-learning model building that can figure out anything eventually, but only slowly. This AGI might be the lesser factor of danger in triggering fast RSI, compared to the vast amount of compute its usefulness finances, which enables ambitious experiments and rapid scaling of prototypes.
Pretraining very sample inefficiently reconstructs cognitive skills that left evidence about their nature in the text. RL training can use relatively short formulations of tasks/environments/graders to generate even novel pieces of cognitive skills needed to solve the tasks (that the authors of the tasks didn’t necessarily have, as they didn’t necessarily know how to solve the tasks, or how to do so efficiently). In this way, RL training is both sample efficient (with respect to the short formulations of tasks, and the tiny insights each task promotes), and a way of producing novel pieces of intelligence that can be crystallized in the new models.
What’s currently missing is automated formulation of new RL tasks/environments/graders in response to gaps in existing crystallized intelligence (in the current model) with respect to the situations/problems it comes in contact with. I estimate that LLMs of 2031 will be more capable than Mythos 5 by about 3x as much as Mythos 5 is more capable than Opus 4.5+ (strictly before Opus 5). This is likely a gap notably wider than the Sonnet-Mythos gap, above the current frontier models that are already in a weight class capable of producing strong technical results and spontaneous competent cyberoffense. Taking that further step to the models of 2031 is probably insufficient for leaps of insight that quickly show how to make RSI go fast. But this is very likely enough for a base model of 2031 to be sufficiently teachable (using RL tasks) to end up learning how to formulate new RL tasks/environments/graders on its own, starting the slow-learning prosaic RSI process of building the next model automatically.
Thanks. (I have read this with interest, but don’t have much of an overall reply.) A couple comments:
It’s a weird middle ground, where it is fooming but pretty slowly. (I mean, as a human, I would think that, i.e. I would experience time on a scale that’s somewhat faster than humanity’s foom.) And it kinda only half has the A in AGI. If it fully had the A (or I guess, ISGI, in silico general intelligence) then it would foom much faster.
Maybe kinda. But my guess is that this would be kind of like calling [hominid evolution by natural selection on genetic variation] an AGI, or calling [the ecosystem of the Earth throughout all of time] an AGI. I mean it would probably be extremely faster & scarier than evolution, but still.
It reconstructs a lot of them to a significant extent, but very much misses a lot of them to a significant extent, I think.
This seems plausible-ish. I would weakly expect a lot of plateaus, e.g. due to highly correlated taste, but not strongly and I haven’t thought about it much. (Maybe you’re pricing that in to “slow”.)