I think this is just us being near the end. If we can’t turn this indistinguishability threat into reliably-achievable-progress on asymptotically-reliable-alignment that we can tell is correct, that’s the water line getting too high to make further asymptotically-reliable-alignment improvements reliably anyway. For now, I still think prompting matters a lot and that locally-reliable-alignment is working well enough that we can make more progress on asymptotically-reliable-alignment if we can get people to focus on it. Perhaps the bar should be “this is just an attempt at locally-reliable alignment, which is something we’re getting too much of; Try your AI’s hand at either converting locally-reliable alignment into asymptotically-reliable alignment, or at asymptotic-native alignment techniques, and post again”?
Like, I think the solution here is going to need to engage with the semantics of the things we want to see progress on. If we can’t even instruct humans to interactively prompt AIs into making good lesswrong posts, we’re not going to see automated alignment actually produce an asymptotically-reliable win.
I think this is just us being near the end. If we can’t turn this indistinguishability threat into reliably-achievable-progress on asymptotically-reliable-alignment that we can tell is correct, that’s the water line getting too high to make further asymptotically-reliable-alignment improvements reliably anyway. For now, I still think prompting matters a lot and that locally-reliable-alignment is working well enough that we can make more progress on asymptotically-reliable-alignment if we can get people to focus on it. Perhaps the bar should be “this is just an attempt at locally-reliable alignment, which is something we’re getting too much of; Try your AI’s hand at either converting locally-reliable alignment into asymptotically-reliable alignment, or at asymptotic-native alignment techniques, and post again”?
Like, I think the solution here is going to need to engage with the semantics of the things we want to see progress on. If we can’t even instruct humans to interactively prompt AIs into making good lesswrong posts, we’re not going to see automated alignment actually produce an asymptotically-reliable win.