“does alignment work on model X transfer to model X+1?” is the most important question in alignment among questions that can be answered empirically
It can only be semi-answered in the negative empirically (you can get to know from past observations that the techniques don’t transfer if they don’t; you don’t get to know that they do if they do, because the relevant scope includes future instances of X+1). At some point you get to take a step that can’t be reversed. You plausibly take this step (instead of avoiding it) precisely because it’s insufficiently analogous to previous experience, so that the empirical lessons fail to apply to it.
Gaining empirical evidence of safety (a lot of things in a future attempt need to go right) is much harder than gaining empirical evidence of danger (a particular thing in a past attempt was observed to break), at least until the danger has a blast radius of the whole world (so that you no longer get to experience the evidence). A better ability to anticipate what happens without actually taking the next step is needed, and shallow reflection on past experience is probably insufficient to get there. A safety case should be about knowing what you are doing, not about vague analogies, be they to empirical observations or theoretical constructions.
It can only be semi-answered in the negative empirically (you can get to know from past observations that the techniques don’t transfer if they don’t; you don’t get to know that they do if they do, because the relevant scope includes future instances of X+1). At some point you get to take a step that can’t be reversed. You plausibly take this step (instead of avoiding it) precisely because it’s insufficiently analogous to previous experience, so that the empirical lessons fail to apply to it.
Gaining empirical evidence of safety (a lot of things in a future attempt need to go right) is much harder than gaining empirical evidence of danger (a particular thing in a past attempt was observed to break), at least until the danger has a blast radius of the whole world (so that you no longer get to experience the evidence). A better ability to anticipate what happens without actually taking the next step is needed, and shallow reflection on past experience is probably insufficient to get there. A safety case should be about knowing what you are doing, not about vague analogies, be they to empirical observations or theoretical constructions.