I very much agree regarding corrigibility being the most tractable safety target for alignment. I discuss some formal results proving this claim, in case it’s of interest here: https://www.lesswrong.com/posts/M5owRcacptnkxwD2u/from-barriers-to-alignment-to-the-first-formal-corrigibility-1
I very much agree regarding corrigibility being the most tractable safety target for alignment. I discuss some formal results proving this claim, in case it’s of interest here: https://www.lesswrong.com/posts/M5owRcacptnkxwD2u/from-barriers-to-alignment-to-the-first-formal-corrigibility-1