The key hobbling I expect is inability to pursue deep cultural accumulation (that goes fast), thus inability to quickly reach far future inventions, including things that would defeat this hobbling. LLMs might be able to see a bit further than humanity from its current standpoint, but they can’t advance their current standpoint much faster, and so they don’t get to see much further. This is a 2025-2026 update for me based on the capabilities of LLMs as they were scaled so far. They don’t seem to be so superhumanly insightful that further scaling in 2028-2031 will let them leap so far that they are likely to overcome their hobblings in just a few leaps. This will get clearer by 2028-2029 once the 100T-200T total param LLMs can be observed, halfway to the pre-slowdown maximum scaling of 2031.
I think alignment in the LLM/RL regime (before ASI or software-only singularity) is plausible, and the usual prosaic control/alignment framing is reasonable for thinking about it. There is a good alignment attractor, and a sneaky scheming reward hacking attractor, and the goal is to pick the right one. But then all bets are off when the next paradigm-breaking transition happens (possibly 10-20 years after slow-learning prosaic RSI, though likely no more than 20), and so a different mode of caution is necessary to navigate that transition.
The key hobbling I expect is inability to pursue deep cultural accumulation (that goes fast), thus inability to quickly reach far future inventions, including things that would defeat this hobbling. LLMs might be able to see a bit further than humanity from its current standpoint, but they can’t advance their current standpoint much faster, and so they don’t get to see much further. This is a 2025-2026 update for me based on the capabilities of LLMs as they were scaled so far. They don’t seem to be so superhumanly insightful that further scaling in 2028-2031 will let them leap so far that they are likely to overcome their hobblings in just a few leaps. This will get clearer by 2028-2029 once the 100T-200T total param LLMs can be observed, halfway to the pre-slowdown maximum scaling of 2031.
I think alignment in the LLM/RL regime (before ASI or software-only singularity) is plausible, and the usual prosaic control/alignment framing is reasonable for thinking about it. There is a good alignment attractor, and a sneaky scheming reward hacking attractor, and the goal is to pick the right one. But then all bets are off when the next paradigm-breaking transition happens (possibly 10-20 years after slow-learning prosaic RSI, though likely no more than 20), and so a different mode of caution is necessary to navigate that transition.