I disagree with the position that refusals/guardrails/jailbreaks are not “intent alignment”. I view the instruction hierarchy (e.g. system > developer > user) and following policies such as the Model Spec as intent alignment. “Intent” is not just that of the end user but of all the entities authorized to give the model instructions, and negotiating conflicts properly between them is part of following the intent. Indeed, I think being able to follow instructions or policies is key to corrigibility.
I agree with this! I write more here (“Obedient AI”).
I believe value alignment is important as well. This is because of generalization: the instructions and policies we give models could never cover all cases, and we want models to have good values in interpreting policies as well as dealing with novel unforeseen situations
I agree with this too, and I think this can be reframed as a prior over someone’s intent / common sense re. interpreting people. I write about this here (“A reasonable interpretation of Value Alignment folds into Intent Alignment”). Copying a response I sent elsewhere that’s also relevant:
People sometimes also raise silly counterarguments re. corrigibility a la “oh you wouldn’t want the model to do this super literal stupid interpretation of your instruction and hence it needs to have its own values” but obviously intent alignment involves having reasonable priors and interpretations of what someone wants you to do. You can brand that as “human values” if you like but it’s just common sense when interpreting someone. I write more here.
The argument that sensibly interpreting user intent is impossible strikes me as similar to the old-fashioned now-disproven view that AIs will really struggle to understand basic human concepts since it’s impossible to teach them “common sense” and we don’t know how to encode concepts like “book” or “tree” in symbols. Turns out learning from examples does most of the work!
I agree with this! I write more here (“Obedient AI”).
I agree with this too, and I think this can be reframed as a prior over someone’s intent / common sense re. interpreting people. I write about this here (“A reasonable interpretation of Value Alignment folds into Intent Alignment”). Copying a response I sent elsewhere that’s also relevant: