But if locally robust alignment makes faster progress on capabilities without being a version of locally robust alignment that can become asymptotically robust alignment and give us a win node, then it’s not only not as obvious, but might well be primarily made of backfire. I don’t believe the most extreme version of this claim, but I used to, I know people who do. The “people going into capabilities” is also including “local alignment that doesn’t generalize to asymptotic alignment”.
The naive way to get asymptotic, and possibly the only way, is to have locally robust alignment that is reliably always ahead of impactfulness of deployment. But if impactfulness of deployment is always operating at 110% of local robustness, that seems like a recipe for disaster. So the question is ultimately one about how to stay in the local validity window indefinitely, even as that local window gets harder to be sure we’ve maintained.
Also there’s the whole thing about how alignment is fundamentally about figuring out what to ask for, and being robust about that is a confusing question at best. Most of MATS seems too focused on the (very hard!) problems of achieving local-alignment-good-enough-to-keep-going to spend the necessary effort to succeed at the really hard parts later (which is not obviously a mistake on MATS’s part, but might well be).
It’s much easier to get reliable robustness for local alignment when you don’t have to solve “what should a thing overwhelmingly smarter than us do?” yet, but if there’s a hole coming where a thing moderately smarter than us has the ability to confuse us into choosing a plan that seems good to it but which is not what we want, we may fall off the manifold we would have wanted to be on if we had been able to see things coming without getting confused.
But if locally robust alignment makes faster progress on capabilities without being a version of locally robust alignment that can become asymptotically robust alignment and give us a win node, then it’s not only not as obvious, but might well be primarily made of backfire. I don’t believe the most extreme version of this claim, but I used to, I know people who do. The “people going into capabilities” is also including “local alignment that doesn’t generalize to asymptotic alignment”.
The naive way to get asymptotic, and possibly the only way, is to have locally robust alignment that is reliably always ahead of impactfulness of deployment. But if impactfulness of deployment is always operating at 110% of local robustness, that seems like a recipe for disaster. So the question is ultimately one about how to stay in the local validity window indefinitely, even as that local window gets harder to be sure we’ve maintained.
Also there’s the whole thing about how alignment is fundamentally about figuring out what to ask for, and being robust about that is a confusing question at best. Most of MATS seems too focused on the (very hard!) problems of achieving local-alignment-good-enough-to-keep-going to spend the necessary effort to succeed at the really hard parts later (which is not obviously a mistake on MATS’s part, but might well be).
It’s much easier to get reliable robustness for local alignment when you don’t have to solve “what should a thing overwhelmingly smarter than us do?” yet, but if there’s a hole coming where a thing moderately smarter than us has the ability to confuse us into choosing a plan that seems good to it but which is not what we want, we may fall off the manifold we would have wanted to be on if we had been able to see things coming without getting confused.