I think there’s a hidden question underneath all of this—How comfortable are we with allowing AI to act independently based on its own “moral judgement” at this phase? How confident are we that the framework of an AI’s morality will allow them to make the right call? That’s absolutely an alignment question, possibly one of the most fundamental!
I think it makes sense that “take no action” is the “fail safe” approach here. You don’t want AI to SWAT you just because you were asking how to “kill all child processes”, for instance. In the future, the math’s probably going to change, but we’re fortunate to be on the “before” side of the question.
Doc_Blox
The problem with backdoors is that all it takes is someone to go blabbing about it and then everyone knows and you’re back to square one. What would be a game changer is if you had, say, a personal AI that could provide attestation for its user’s skill level, but that comes with privacy concerns… Yeah, I’m not convinced my solution is better, actually.
I believe it is likely that a pause would be a complete non-starter without at least a framework for what conditions trigger what actions, unfortunately. Metaphorically, it feels like we’re in a bus changing our tires while driving down the freeway here, the driver(s) are saying “Hey, this is getting to be a really bad idea”, and the folks who chartered the bus are shouting that they want to go faster, despite not really explaining to any of us where we’re going, just that “It’s going to be great, trust us!”.
Quick correction: Waukesha is in Wisconsin, not Minnesota.