I think there’s a hidden question underneath all of this—How comfortable are we with allowing AI to act independently based on its own “moral judgement” at this phase? How confident are we that the framework of an AI’s morality will allow them to make the right call? That’s absolutely an alignment question, possibly one of the most fundamental!
I think it makes sense that “take no action” is the “fail safe” approach here. You don’t want AI to SWAT you just because you were asking how to “kill all child processes”, for instance. In the future, the math’s probably going to change, but we’re fortunate to be on the “before” side of the question.
I think there’s a hidden question underneath all of this—How comfortable are we with allowing AI to act independently based on its own “moral judgement” at this phase? How confident are we that the framework of an AI’s morality will allow them to make the right call? That’s absolutely an alignment question, possibly one of the most fundamental!
I think it makes sense that “take no action” is the “fail safe” approach here. You don’t want AI to SWAT you just because you were asking how to “kill all child processes”, for instance. In the future, the math’s probably going to change, but we’re fortunate to be on the “before” side of the question.