Is there something that is built-in to your definition of alignment that leads you to believe otherwise? In other words, what do you think I’m missing about the definition of alignment for this to not be the case?
(I believe AI systems differ from traditional systems in that AI systems may be intelligently optimizing for a goal that is different from what their creators initially built them to optimize for. From my understanding, (inner) alignment is about fixing that mismatch. I give this example with the plane to point out that even if we solve alignment, we still need AI safety research beyond alignment research to ensure that such aligned AIs are actually safe for humans.)
I think that if you define alignment this way, saying that there can be aligned AIs that kill people should be met with “of course they can, because that is true of just about every machine with any physical capability at all”. If a car counts as an aligned system that kills people, then aligned systems that kill people are acceptable.
Assuming we’re talking about non-AI systems, from my understanding of commonly-used definitions of alignment, yes.
Is there something that is built-in to your definition of alignment that leads you to believe otherwise? In other words, what do you think I’m missing about the definition of alignment for this to not be the case?
(I believe AI systems differ from traditional systems in that AI systems may be intelligently optimizing for a goal that is different from what their creators initially built them to optimize for. From my understanding, (inner) alignment is about fixing that mismatch. I give this example with the plane to point out that even if we solve alignment, we still need AI safety research beyond alignment research to ensure that such aligned AIs are actually safe for humans.)
I think that if you define alignment this way, saying that there can be aligned AIs that kill people should be met with “of course they can, because that is true of just about every machine with any physical capability at all”. If a car counts as an aligned system that kills people, then aligned systems that kill people are acceptable.