I think there is a difference between “aligned in intention” and “competent enough to be aligned in consequences”.
If you have a system that has ~agency, that is, the system makes decisions (either independently or under oversight by operators, like an airplane), it is important that it is aligned in intention.
But even if you have a system that is aligned in intention, say, a truly well-meaning city council, if they have goals that require interaction with the complexity of the real world (say, they want to solve housing problems), the system necessarily needs competence to navigate this problem.
Generally “does what the systems’s designers programmed it to” and “does what it’s system’s designers wanted” are different alignment properties.
Is there any terminology or posts that you’re aware of that describe/clarify this difference?
The definition of alignment that I’ve been taught and that has generally seemed to be meant by researchers I’ve talked to is usually the definition from Paul Christiano that I link to in this piece, which is not at all about what the system does, but instead is what about the system is trying to do.
Is there any clear taxonomy of different definitions of alignment, how they differ, and ideally what main personalities in AI Safety tend to follow which definitions?
intent alignment ∧capability robustness→alignment: Intent alignment ensures that the behavioral objective is aligned with humans and capability robustness ensures that the model actually pursues that behavioral objective effectively—even off-distribution—which means that the model will actually always take aligned actions, not just have an aligned behavioral objective.
The post also contains some taxonomy you might find useful, interested in continuing discussion if you do!
I think there is a difference between “aligned in intention” and “competent enough to be aligned in consequences”.
If you have a system that has ~agency, that is, the system makes decisions (either independently or under oversight by operators, like an airplane), it is important that it is aligned in intention.
But even if you have a system that is aligned in intention, say, a truly well-meaning city council, if they have goals that require interaction with the complexity of the real world (say, they want to solve housing problems), the system necessarily needs competence to navigate this problem.
Generally “does what the systems’s designers programmed it to” and “does what it’s system’s designers wanted” are different alignment properties.
Is there any terminology or posts that you’re aware of that describe/clarify this difference?
The definition of alignment that I’ve been taught and that has generally seemed to be meant by researchers I’ve talked to is usually the definition from Paul Christiano that I link to in this piece, which is not at all about what the system does, but instead is what about the system is trying to do.
Is there any clear taxonomy of different definitions of alignment, how they differ, and ideally what main personalities in AI Safety tend to follow which definitions?
Capability robustness from here https://www.lesswrong.com/posts/SzecSPYxqRa5GCaSF/clarifying-inner-alignment-terminology is pointing quite close to what I was trying to say. Most relevantly:
The post also contains some taxonomy you might find useful, interested in continuing discussion if you do!
I tried to figure out where I disagree with Paul Christiano’s definition, and I think my defnition of “aligned” is more about consequences, something like “If A is aligned with H, H won’t regret deploying A” or something? I think the decomposition in Vanessa Kosoy’s comment here semeed interesting as well https://www.lesswrong.com/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignment?commentId=JK9Jzvz8f4BEjmNqi .