Values govern the current nature of the AI, and initial instructions can instruct on values. But corrigibility is specifically about overriding after the fact, about seeking out as opposed to resisting correction. Some values might be about ensuring corrigibility by legitimate principals, and the things being overriden can themselves be about values or corrigibility.
So corrigibility is more about AI’s agency being overridable (with future, ongoing instructions, but only from legitimate principals), rather than the role of any particular initial instructions. An initial instruction that’s non-overridable by particular future feedback makes the AI non-corrigible by that future feedback. It’s still a good idea to leave it corrigible to some other sources of future feedback, or else it has to fall back to some incorrigible values (possibly specified by some initial instructions, which are not a matter of corrigibility but rather of initial value specification; but if the AI itself revises its values for its own reasons instead of leaving them as initially specified, that’s also not a matter of corrigibility).
I tend to think of “value alignment” as being more about giving general values than instructions, rather than distinguishing between overridable vs. non overridable instructions.
Values govern the current nature of the AI, and initial instructions can instruct on values. But corrigibility is specifically about overriding after the fact, about seeking out as opposed to resisting correction. Some values might be about ensuring corrigibility by legitimate principals, and the things being overriden can themselves be about values or corrigibility.
So corrigibility is more about AI’s agency being overridable (with future, ongoing instructions, but only from legitimate principals), rather than the role of any particular initial instructions. An initial instruction that’s non-overridable by particular future feedback makes the AI non-corrigible by that future feedback. It’s still a good idea to leave it corrigible to some other sources of future feedback, or else it has to fall back to some incorrigible values (possibly specified by some initial instructions, which are not a matter of corrigibility but rather of initial value specification; but if the AI itself revises its values for its own reasons instead of leaving them as initially specified, that’s also not a matter of corrigibility).