Whether the simplest function is misaligned (as posited by List of Lethalities #20) is the thing you have to explain!
Is this in fact a crux for you? If you were largely convinced that the simplest functions found by gradient descent in the current paradigm would not remotely approximate human values, to what extent would this shift your odds of the current paradigm getting everyone killed?
I mean, yes. What else could I possibly say? Of course, yes.
In the spirit of not trying to solve the entire alignment problem at once, I find it hard to be too specific about to what extent how my odds would shift without a more specific question. (I think LLMs are doing a pretty good job of knowing and doing what I mean, which implies some form of knowledge of “human values”, but it’s only a natural-language instruction-follower; it’s not supposed to be a sovereign superintelligence, which looks vastly harder and I would rather people not do that for a long time.) Show me the ArXiv paper about inductive biases that I’m supposed to be updating on, and I’ll tell you how much more terrified I am (above my baseline of “already pretty terrified, actually”).
Is this in fact a crux for you? If you were largely convinced that the simplest functions found by gradient descent in the current paradigm would not remotely approximate human values, to what extent would this shift your odds of the current paradigm getting everyone killed?
I mean, yes. What else could I possibly say? Of course, yes.
In the spirit of not trying to solve the entire alignment problem at once, I find it hard to be too specific about to what extent how my odds would shift without a more specific question. (I think LLMs are doing a pretty good job of knowing and doing what I mean, which implies some form of knowledge of “human values”, but it’s only a natural-language instruction-follower; it’s not supposed to be a sovereign superintelligence, which looks vastly harder and I would rather people not do that for a long time.) Show me the ArXiv paper about inductive biases that I’m supposed to be updating on, and I’ll tell you how much more terrified I am (above my baseline of “already pretty terrified, actually”).