Interesting. I was putting everything that could prevent a super-intelligence from taking control in the category of alignment, but the boundary between that and control is fuzzy and subject to different definitions.
Approaches like Internal independent review straddle that line; they are situated outside the LLM itself but serve to prevent minor misalignments from doing damage and growing into egregious misalignments (via memetic spread through memory and goal representations).
Interesting. I was putting everything that could prevent a super-intelligence from taking control in the category of alignment, but the boundary between that and control is fuzzy and subject to different definitions.
Approaches like Internal independent review straddle that line; they are situated outside the LLM itself but serve to prevent minor misalignments from doing damage and growing into egregious misalignments (via memetic spread through memory and goal representations).