I like the distinction between capabilities as a broad training target and alignment as a narrow one. One way to make part of this precise is through shortcut learning: if intended and spurious features are correlated in the training feedback, both can reduce the training loss, so successful in-distribution behavior does not identify which features drive the model’s behavior.
In our recent ICML work on preference optimization, a local linear analysis shows how these spurious parameters are driven by feature means and causal–spurious cross-covariances. When those relationships change at deployment, the learned alignment behavior need not generalize.
I like the distinction between capabilities as a broad training target and alignment as a narrow one. One way to make part of this precise is through shortcut learning: if intended and spurious features are correlated in the training feedback, both can reduce the training loss, so successful in-distribution behavior does not identify which features drive the model’s behavior.
In our recent ICML work on preference optimization, a local linear analysis shows how these spurious parameters are driven by feature means and causal–spurious cross-covariances. When those relationships change at deployment, the learned alignment behavior need not generalize.