What type/capability level of AI are you referring to when you say this? Your separating loss of control from alignment makes me think you’re addressing only AI like current systems, which is not far above human intelligence levels.
The position seems totally reasonable. For superintelligence, alignment will just have to do more of the work against loss of control, which it can even if it’s imperfect.
Interesting. I was putting everything that could prevent a super-intelligence from taking control in the category of alignment, but the boundary between that and control is fuzzy and subject to different definitions.
Approaches like Internal independent review straddle that line; they are situated outside the LLM itself but serve to prevent minor misalignments from doing damage and growing into egregious misalignments (via memetic spread through memory and goal representations).
Yes, I think you’re right that I mainly imagine AI mostly like current systems. This already seems scary enough to pose a large danger via power concentration / internal deployment / gradual disempowerment. And my current sense is that these threats will be the main thing to worry about quite long before we need to seriously worry about loss of control.
But this is a weakly held take / I consider myself not very well informed here.
What type/capability level of AI are you referring to when you say this? Your separating loss of control from alignment makes me think you’re addressing only AI like current systems, which is not far above human intelligence levels.
The position seems totally reasonable. For superintelligence, alignment will just have to do more of the work against loss of control, which it can even if it’s imperfect.
Interesting. I was putting everything that could prevent a super-intelligence from taking control in the category of alignment, but the boundary between that and control is fuzzy and subject to different definitions.
Approaches like Internal independent review straddle that line; they are situated outside the LLM itself but serve to prevent minor misalignments from doing damage and growing into egregious misalignments (via memetic spread through memory and goal representations).
Yes, I think you’re right that I mainly imagine AI mostly like current systems. This already seems scary enough to pose a large danger via power concentration / internal deployment / gradual disempowerment. And my current sense is that these threats will be the main thing to worry about quite long before we need to seriously worry about loss of control.
But this is a weakly held take / I consider myself not very well informed here.