intent alignment ∧capability robustness→alignment: Intent alignment ensures that the behavioral objective is aligned with humans and capability robustness ensures that the model actually pursues that behavioral objective effectively—even off-distribution—which means that the model will actually always take aligned actions, not just have an aligned behavioral objective.
The post also contains some taxonomy you might find useful, interested in continuing discussion if you do!
Capability robustness from here https://www.lesswrong.com/posts/SzecSPYxqRa5GCaSF/clarifying-inner-alignment-terminology is pointing quite close to what I was trying to say. Most relevantly:
The post also contains some taxonomy you might find useful, interested in continuing discussion if you do!
I tried to figure out where I disagree with Paul Christiano’s definition, and I think my defnition of “aligned” is more about consequences, something like “If A is aligned with H, H won’t regret deploying A” or something? I think the decomposition in Vanessa Kosoy’s comment here semeed interesting as well https://www.lesswrong.com/posts/ZeE7EKHTFMBs8eMxn/clarifying-ai-alignment?commentId=JK9Jzvz8f4BEjmNqi .