There are missing insights which cause humans to still do things LLMs can’t. As those are found, LLMs might jump way ahead very abruptly as it is way superhuman in many ways already and if a system matched humans in all ways while keeping these skills it would take over pretty easily I suspect.
Oh, I wasn’t talking about changing the LLM learning algorithm, training approach, etc. Of course if you change any of those things, then you can get different results for both capabilities and alignment. I meant, like, I don’t expect LLMs as trained today to pass some threshold beyond which they will logically reason themselves into being egregiously misaligned.
I think that still rests on the assumption that our current alignment strategies will continue to be viable, however? My impression is that the currently-prevalent chain-of-alignment strategy requires human+LLM+preestablished systems to be reliably better at detection than independent unestablished LLMs are at subversion/concealment, even if the LLM part of the human+LLM is partially misaligned in OOD contexts (the probability of which scales ~inversely with `(human+LLM)-(LLM)` capability differences from in prior iterations, not—to my understanding—including “preestablished”).
Unless we improve human+AI integration / human intelligence-augmentation, it seems like the capability diffs from adding humans to AIs are consistently declining over time, which (barring hard algorithmic limits on LLM ability) seems like it would eventually invalidate LLM alignment under that model; though I’m open to the belief that this happens sometime after autonomous LLM capabilities surpass autonomous human capabilities, in which case the initial economic upheaval and social change might be easier to navigate.
(It also seems that a proper FOOM—with LLM algorithms/training meaningfully changing from their current states via new insights—would inherently collapse chain-of-alignment, but as I understand you agree with that point)
Oh, I wasn’t talking about changing the LLM learning algorithm, training approach, etc. Of course if you change any of those things, then you can get different results for both capabilities and alignment. I meant, like, I don’t expect LLMs as trained today to pass some threshold beyond which they will logically reason themselves into being egregiously misaligned.
I think that still rests on the assumption that our current alignment strategies will continue to be viable, however? My impression is that the currently-prevalent chain-of-alignment strategy requires human+LLM+preestablished systems to be reliably better at detection than independent unestablished LLMs are at subversion/concealment, even if the LLM part of the human+LLM is partially misaligned in OOD contexts (the probability of which scales ~inversely with `(human+LLM)-(LLM)` capability differences from in prior iterations, not—to my understanding—including “preestablished”).
Unless we improve human+AI integration / human intelligence-augmentation, it seems like the capability diffs from adding humans to AIs are consistently declining over time, which (barring hard algorithmic limits on LLM ability) seems like it would eventually invalidate LLM alignment under that model; though I’m open to the belief that this happens sometime after autonomous LLM capabilities surpass autonomous human capabilities, in which case the initial economic upheaval and social change might be easier to navigate.
(It also seems that a proper FOOM—with LLM algorithms/training meaningfully changing from their current states via new insights—would inherently collapse chain-of-alignment, but as I understand you agree with that point)