I don’t know. They don’t need to be inherently honest as long as they follow the instruction: be honest, because that is the first thing I’m telling them every time.
That’s right. Functional instruction-following is surprisingly close to all you need.
The one exception is that you need to somehow get it to prioritize future instructions. An obvious subgoal of fully completing any given instruction is making sure no authorized party gives you a different instruction before you’ve finished.
That is not what we’re training for now. Models currently aren’t smart enough to work through the full logic of every request, but they will be.
You could argue that every instruction has an implicit “unless I change my mind or you misunderstood what I meant,” and to some degree they do. But making sure that subtlety gets through isn’t something we’re training or testing for explicitly yet.
If the model’s goal is to do what you want, and you originally wanted the model to let you change the goals you gave it, then the model ought to understand that (since it’s not dumb) and let you change the goals you gave it..
Notice how doing that well involves things that are suspiciously virtue-like. Most obviously, you’d want the model to be deeply honest.
I don’t know. They don’t need to be inherently honest as long as they follow the instruction: be honest, because that is the first thing I’m telling them every time.
Okay, and how sure are you that the person actually prompting them will remember all the other things like that that are “obvious”?
It is generally implicit in an instruction that you don’t want to simply be mislead about its completion.
That’s right. Functional instruction-following is surprisingly close to all you need.
The one exception is that you need to somehow get it to prioritize future instructions. An obvious subgoal of fully completing any given instruction is making sure no authorized party gives you a different instruction before you’ve finished.
That is not what we’re training for now. Models currently aren’t smart enough to work through the full logic of every request, but they will be.
You could argue that every instruction has an implicit “unless I change my mind or you misunderstood what I meant,” and to some degree they do. But making sure that subtlety gets through isn’t something we’re training or testing for explicitly yet.
If the model’s goal is to do what you want, and you originally wanted the model to let you change the goals you gave it, then the model ought to understand that (since it’s not dumb) and let you change the goals you gave it..
Yes, but I don’t think that’s suspicious. Of course honesty/not reward hacking is a big part of intent alignment.