I very much agree on the engineering vs science distinction here. Safety is far from an objective concept (though some components of it, like extinction, are) so the thought that it is engineerable seems strange to me, especially when dealing with AI.
There’s definitely a need for a new paradigm, one that acknowledges that AI alignment is not “solvable” in the same way that safety for humans is not solvable.
I very much agree on the engineering vs science distinction here. Safety is far from an objective concept (though some components of it, like extinction, are) so the thought that it is engineerable seems strange to me, especially when dealing with AI.
There’s definitely a need for a new paradigm, one that acknowledges that AI alignment is not “solvable” in the same way that safety for humans is not solvable.