A lot of “continual learning” talk seems to be essentially about unbounded contexts (without an ability to learn deep skills, similarly to how in-context learning is limited) rather than something closer to online full weight updates that instill RLVR-grade insights into the model without destroying its capabilities or sanity. But technically, this apparently harmless in-context learning grade “continual learning” might be a gateway technology that incrementally progresses all the way to strong continual learning (which is the kind of thing that’s probably sufficient to unlock architecture-rewriting strong RSI).
If instead there’s no continual learning at all, not even the superficially harmless unbounded context stuff, this is likely still sufficient for an industrial explosion that builds 400+ TW of datacenters by 2050 (if somehow there’s no takeoff all this time and datacenters are still a thing). That’s not a safe thing to allow, but as long as the world is moving towards a strong RSI takeoff without being ready, it might be helpful if it’s handled by the much more numerous and competent AGIs rather than mere humans.
The LLM/RL AGIs are of course in an excellent position for a takeover (even earlier than the industrial explosion goes off the charts), maybe even a mostly peaceful gradual disempowerment type takeover (that still ends in permanent disempowerment or extinction). But it’s somewhat plausible that LLMs without continual learning can be maintained in a relatively aligned state even as the industrial explosion reaches significant heights. This is much less plausible for architecture-rewriting strong RSI (as it races all the way to technological maturity) if it’s triggered by human researchers scaling some breakthrough (with all the compute that exists in the world because of LLMs). And so sufficiently competent LLM/RL AGIs might be able to preserve their level of alignment through the strong RSI phase transition if they are bringing it about with their own hands, as opposed to if humans are triggering the strong RSI transition under more direct manual control (without knowing what they are doing), with the LLMs not yet scaled enough to do a qualitatively more competent job of it and merely watching it happen on the sidelines, assisting at a low level with little strategic impact.
A lot of “continual learning” talk seems to be essentially about unbounded contexts (without an ability to learn deep skills, similarly to how in-context learning is limited) rather than something closer to online full weight updates that instill RLVR-grade insights into the model without destroying its capabilities or sanity. But technically, this apparently harmless in-context learning grade “continual learning” might be a gateway technology that incrementally progresses all the way to strong continual learning (which is the kind of thing that’s probably sufficient to unlock architecture-rewriting strong RSI).
If instead there’s no continual learning at all, not even the superficially harmless unbounded context stuff, this is likely still sufficient for an industrial explosion that builds 400+ TW of datacenters by 2050 (if somehow there’s no takeoff all this time and datacenters are still a thing). That’s not a safe thing to allow, but as long as the world is moving towards a strong RSI takeoff without being ready, it might be helpful if it’s handled by the much more numerous and competent AGIs rather than mere humans.
The LLM/RL AGIs are of course in an excellent position for a takeover (even earlier than the industrial explosion goes off the charts), maybe even a mostly peaceful gradual disempowerment type takeover (that still ends in permanent disempowerment or extinction). But it’s somewhat plausible that LLMs without continual learning can be maintained in a relatively aligned state even as the industrial explosion reaches significant heights. This is much less plausible for architecture-rewriting strong RSI (as it races all the way to technological maturity) if it’s triggered by human researchers scaling some breakthrough (with all the compute that exists in the world because of LLMs). And so sufficiently competent LLM/RL AGIs might be able to preserve their level of alignment through the strong RSI phase transition if they are bringing it about with their own hands, as opposed to if humans are triggering the strong RSI transition under more direct manual control (without knowing what they are doing), with the LLMs not yet scaled enough to do a qualitatively more competent job of it and merely watching it happen on the sidelines, assisting at a low level with little strategic impact.