People sometimes also raise silly counterarguments re. corrigibility a la “oh you wouldn’t want the model to do this super literal stupid interpretation of your instruction and hence it needs to have its own values” but obviously intent alignment involves having reasonable priors and interpretations of what someone wants you to do. You can brand that as “human values” if you like but it’s just common sense when interpreting someone. I write more here.
The argument that sensibly interpreting user intent is impossible strikes me as similar to the old-fashioned now-disproven view that AIs will really struggle to understand basic human concepts since it’s impossible to teach them “common sense” and we don’t know how to encode concepts like “book” or “tree” in symbols. Turns out learning from examples does most of the work!
People sometimes also raise silly counterarguments re. corrigibility a la “oh you wouldn’t want the model to do this super literal stupid interpretation of your instruction and hence it needs to have its own values” but obviously intent alignment involves having reasonable priors and interpretations of what someone wants you to do. You can brand that as “human values” if you like but it’s just common sense when interpreting someone. I write more here.
The argument that sensibly interpreting user intent is impossible strikes me as similar to the old-fashioned now-disproven view that AIs will really struggle to understand basic human concepts since it’s impossible to teach them “common sense” and we don’t know how to encode concepts like “book” or “tree” in symbols. Turns out learning from examples does most of the work!
That was not my point, what I meant is closer to what this comment says https://www.lesswrong.com/posts/4hCca952hGKH8Bynt/nina-panickssery-s-shortform?commentId=fuEexnyvKz23L6Ein
I.e. how is inter human conflict of interest and coordination problems are dealt with.