Is there any terminology or posts that you’re aware of that describe/clarify this difference?
The definition of alignment that I’ve been taught and that has generally seemed to be meant by researchers I’ve talked to is usually the definition from Paul Christiano that I link to in this piece, which is not at all about what the system does, but instead is what about the system is trying to do.
Is there any clear taxonomy of different definitions of alignment, how they differ, and ideally what main personalities in AI Safety tend to follow which definitions?
intent alignment ∧capability robustness→alignment: Intent alignment ensures that the behavioral objective is aligned with humans and capability robustness ensures that the model actually pursues that behavioral objective effectively—even off-distribution—which means that the model will actually always take aligned actions, not just have an aligned behavioral objective.
The post also contains some taxonomy you might find useful, interested in continuing discussion if you do!
Is there any terminology or posts that you’re aware of that describe/clarify this difference?
The definition of alignment that I’ve been taught and that has generally seemed to be meant by researchers I’ve talked to is usually the definition from Paul Christiano that I link to in this piece, which is not at all about what the system does, but instead is what about the system is trying to do.
Is there any clear taxonomy of different definitions of alignment, how they differ, and ideally what main personalities in AI Safety tend to follow which definitions?
Capability robustness from here https://www.lesswrong.com/posts/SzecSPYxqRa5GCaSF/clarifying-inner-alignment-terminology is pointing quite close to what I was trying to say. Most relevantly:
The post also contains some taxonomy you might find useful, interested in continuing discussion if you do!
I tried to figure out where I disagree with Paul Christiano’s definition, and I think my defnition of “aligned” is more about consequences, something like “If A is aligned with H, H won’t regret deploying A” or something? I think the decomposition in Vanessa Kosoy’s comment here semeed interesting as well https://www.lesswrong.com/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignment?commentId=JK9Jzvz8f4BEjmNqi .