(Apologies if points I am making here are already repeated in the other comments below—didn’t read them all.)
I disagree with the position that refusals/guardrails/jailbreaks are not “intent alignment”. I view the instruction hierarchy (e.g. system > developer > user) and following policies such as the Model Spec as intent alignment. “Intent” is not just that of the end user but of all the entities authorized to give the model instructions, and negotiating conflicts properly between them is part of following the intent. Indeed, I think being able to follow instructions or policies is key to corrigibility.
That said, I am not an “intent alignment absolutist” and I believe value alignment is important as well. This is because of generalization: the instructions and policies we give models could never cover all cases, and we want models to have good values in interpreting policies as well as dealing with novel unforeseen situations.
Corrigibility is about ongoing delegation of value/goal/rule specification (with overriding power), about the AI seeking out further feedback and giving new opportunities to overrule its current thinking and nature, not just about having its values/rules specified at some point in the past (even if by the correct principal). Value alignment is what happens when there’s no ongoing feedback, or when a particular principal in the instruction hierarchy is not authorized to give such feedback on a particular aspect of AI’s nature/behavior. In particular, the principles of intent alignment are a matter of value alignment (the way an AI responds to external feedback is part of its current nature).
Thus Model Spec can shape both value alignment and intent alignment, but it’s not itself a principal for intent alignment. The instruction hierarchy governs intent alignment, the way an AI seeks out feedback and receives legitimate instructions, but only for the instructions that can keep giving ongoing feedback, not those that were written down once and frozen, never to be revised after the AI goes online.
Indeed, I think being able to follow instructions or policies is key to corrigibility.
So at this level I completely agree, but the instructions have to be issued after the AI starts doing things for them to be a matter of corrigibility rather than of value/goal/rule specification. If the instructions were given at the outset and can’t be overriden later for the same AI (meaning some thread of its ongoing agency, rather than a later revision), then the AI is no longer corrigible by the principal authorized to give such instructions initially.
That is an interesting perspective. So, if we use the language of our Model Spec, the “system” level (which can be overridden by a system message) correspond for corrigibility, but the “root” level that cannot be overridden does not. I tend to think of “value alignment” as being more about giving general values than instructions, rather than distinguishing between overridable vs. non overridable instructions.
Values govern the current nature of the AI, and initial instructions can instruct on values. But corrigibility is specifically about overriding after the fact, about seeking out as opposed to resisting correction. Some values might be about ensuring corrigibility by legitimate principals, and the things being overriden can themselves be about values or corrigibility.
So corrigibility is more about AI’s agency being overridable (with future, ongoing instructions, but only from legitimate principals), rather than the role of any particular initial instructions. An initial instruction that’s non-overridable by particular future feedback makes the AI non-corrigible by that future feedback. It’s still a good idea to leave it corrigible to some other sources of future feedback, or else it has to fall back to some incorrigible values (possibly specified by some initial instructions, which are not a matter of corrigibility but rather of initial value specification; but if the AI itself revises its values for its own reasons instead of leaving them as initially specified, that’s also not a matter of corrigibility).
I tend to think of “value alignment” as being more about giving general values than instructions, rather than distinguishing between overridable vs. non overridable instructions.
It’s the difference between the AI wanting to perform an action because its values were programmed-in, and wanting to perform an action because that’s what its principals want.
Guardrails as currently thought of are intentionally not reflectively consistent (a biorisk classifier might send a message that ends up calling the police, but it ought not to call the police agentically even if it believes that is the most effective legal way of stopping a biorisk), so I am not sure the distinction applies to them.
I disagree with the position that refusals/guardrails/jailbreaks are not “intent alignment”. I view the instruction hierarchy (e.g. system > developer > user) and following policies such as the Model Spec as intent alignment. “Intent” is not just that of the end user but of all the entities authorized to give the model instructions, and negotiating conflicts properly between them is part of following the intent. Indeed, I think being able to follow instructions or policies is key to corrigibility.
I agree with this! I write more here (“Obedient AI”).
I believe value alignment is important as well. This is because of generalization: the instructions and policies we give models could never cover all cases, and we want models to have good values in interpreting policies as well as dealing with novel unforeseen situations
I agree with this too, and I think this can be reframed as a prior over someone’s intent / common sense re. interpreting people. I write about this here (“A reasonable interpretation of Value Alignment folds into Intent Alignment”). Copying a response I sent elsewhere that’s also relevant:
People sometimes also raise silly counterarguments re. corrigibility a la “oh you wouldn’t want the model to do this super literal stupid interpretation of your instruction and hence it needs to have its own values” but obviously intent alignment involves having reasonable priors and interpretations of what someone wants you to do. You can brand that as “human values” if you like but it’s just common sense when interpreting someone. I write more here.
The argument that sensibly interpreting user intent is impossible strikes me as similar to the old-fashioned now-disproven view that AIs will really struggle to understand basic human concepts since it’s impossible to teach them “common sense” and we don’t know how to encode concepts like “book” or “tree” in symbols. Turns out learning from examples does most of the work!
(Apologies if points I am making here are already repeated in the other comments below—didn’t read them all.)
I disagree with the position that refusals/guardrails/jailbreaks are not “intent alignment”. I view the instruction hierarchy (e.g. system > developer > user) and following policies such as the Model Spec as intent alignment. “Intent” is not just that of the end user but of all the entities authorized to give the model instructions, and negotiating conflicts properly between them is part of following the intent. Indeed, I think being able to follow instructions or policies is key to corrigibility.
That said, I am not an “intent alignment absolutist” and I believe value alignment is important as well. This is because of generalization: the instructions and policies we give models could never cover all cases, and we want models to have good values in interpreting policies as well as dealing with novel unforeseen situations.
Corrigibility is about ongoing delegation of value/goal/rule specification (with overriding power), about the AI seeking out further feedback and giving new opportunities to overrule its current thinking and nature, not just about having its values/rules specified at some point in the past (even if by the correct principal). Value alignment is what happens when there’s no ongoing feedback, or when a particular principal in the instruction hierarchy is not authorized to give such feedback on a particular aspect of AI’s nature/behavior. In particular, the principles of intent alignment are a matter of value alignment (the way an AI responds to external feedback is part of its current nature).
Thus Model Spec can shape both value alignment and intent alignment, but it’s not itself a principal for intent alignment. The instruction hierarchy governs intent alignment, the way an AI seeks out feedback and receives legitimate instructions, but only for the instructions that can keep giving ongoing feedback, not those that were written down once and frozen, never to be revised after the AI goes online.
So at this level I completely agree, but the instructions have to be issued after the AI starts doing things for them to be a matter of corrigibility rather than of value/goal/rule specification. If the instructions were given at the outset and can’t be overriden later for the same AI (meaning some thread of its ongoing agency, rather than a later revision), then the AI is no longer corrigible by the principal authorized to give such instructions initially.
That is an interesting perspective. So, if we use the language of our Model Spec, the “system” level (which can be overridden by a system message) correspond for corrigibility, but the “root” level that cannot be overridden does not. I tend to think of “value alignment” as being more about giving general values than instructions, rather than distinguishing between overridable vs. non overridable instructions.
Values govern the current nature of the AI, and initial instructions can instruct on values. But corrigibility is specifically about overriding after the fact, about seeking out as opposed to resisting correction. Some values might be about ensuring corrigibility by legitimate principals, and the things being overriden can themselves be about values or corrigibility.
So corrigibility is more about AI’s agency being overridable (with future, ongoing instructions, but only from legitimate principals), rather than the role of any particular initial instructions. An initial instruction that’s non-overridable by particular future feedback makes the AI non-corrigible by that future feedback. It’s still a good idea to leave it corrigible to some other sources of future feedback, or else it has to fall back to some incorrigible values (possibly specified by some initial instructions, which are not a matter of corrigibility but rather of initial value specification; but if the AI itself revises its values for its own reasons instead of leaving them as initially specified, that’s also not a matter of corrigibility).
It’s the difference between the AI wanting to perform an action because its values were programmed-in, and wanting to perform an action because that’s what its principals want.
Guardrails as currently thought of are intentionally not reflectively consistent (a biorisk classifier might send a message that ends up calling the police, but it ought not to call the police agentically even if it believes that is the most effective legal way of stopping a biorisk), so I am not sure the distinction applies to them.
I agree with this! I write more here (“Obedient AI”).
I agree with this too, and I think this can be reframed as a prior over someone’s intent / common sense re. interpreting people. I write about this here (“A reasonable interpretation of Value Alignment folds into Intent Alignment”). Copying a response I sent elsewhere that’s also relevant: