Curated. This piece was extremely easy to read and I like how it clearly points out that some kinds of misalignment are more likely given certain kinds of training setups. I also like taxonomies of possible kinds of misalignment in general. It reminds me of this section of the appendix of Ryan’s Current AIs seem pretty misaligned to me. I’d like to see more taxonomies of possible kinds of misalignment and when and why we can expect them to show up. This seemed straightforward at least in retrospect, but I am glad to have it written up somewhere.
Curated. This piece was extremely easy to read and I like how it clearly points out that some kinds of misalignment are more likely given certain kinds of training setups. I also like taxonomies of possible kinds of misalignment in general. It reminds me of this section of the appendix of Ryan’s Current AIs seem pretty misaligned to me. I’d like to see more taxonomies of possible kinds of misalignment and when and why we can expect them to show up. This seemed straightforward at least in retrospect, but I am glad to have it written up somewhere.