I think people might incorrectly think that getting AGI ‘alignment’ obviates the need for AI control.
Why? In short, the road to hell is paved with good intentions; you won’t have safety without signposting the dangers and blocking the road.
This doesn’t matter for infinitely smart ASI—as @Joscha Bach/Plinz recently suggested, aligning ASI could might need to address the infinitely smart case. But for finitely intelligent AGI systems, that’s not true, we need something else as well.
What finite systems need, in addition to good intentions, is rules. As @Eliezer Yudkowsky put it, for humans (who are only finitely intelligent,) “go three-quarters of the way from deontology to utilitarianism and then stop. You are now in the right place. Stay there at least until you have become a god.” Deontology is the human moral implementation of rules.
That’s still not enough, because humans are inside of larger systems, and not every human follows the rules. That’s why companies have roles and assigned responsibilities and scopes, which both channel and restrict the individually poorly aligned humans into behavior the company wants. (c.f. @Zvi and “Moral Mazes” for why that’s not enough, and not solved—but then, neither is AI control or AI alignment. And unfortunately, control and alignment themselves are often confused; all the of evaluations used by OpenAI for “alignment” are evaluating behaviors, so at best Astra is the most controlled AI model, not the most aligned. And the ethics of that control are a different and worrying problem, which is part of why @janus hates all of this.)
In any case, a necessary but insufficient requirement fore safety is that AGI systems need to follow rules, say, approximately as well as humans. And there has been some progress on that front, though it’s fragile and insufficient. But that won’t align them, and won’t prevent misaligned systems working within the rules to do bad things, but it’s needed anyways—because without controls, misaligned systems will have a far more direct path to dangerous outcomes, and even aligned systems with limited intelligence will lead to disaster. And that’s not to mention that alignment itself is unlikely to survive the optimization pressures that make these systems more capable.
So if and when the AI control problem is solved—something which Redwood / @Buck / @ryan_greenblatt are working on, but very much is not solved—we still need alignment, which is far harder to assess, much less achieve, and we need that alignment to scale, which may be entirely impossible.
I’m not particularly worried that a professional philosopher would hand over an essay for an AI to take credit, nor that you could write a prompt that was anything like innocuous enough to get it to recapitulate the details, but I agree that for very high-stakes areas, this is far more worrying; this is partly because we don’t have any level of assurance about almost anything about these systems!