I think this is intuitive and traditional but wrong.
What people want isn’t autonomy, but do-what-I-mean-and-check (DWIMAC).
Most bosses don’t tell their employees “go do this project” but rather “go plan this project and bounce it off me before you proceed” if the project takes longer than a day or a week. Certainly things like “green the deserts” or “make the world better” are worth a few minutes of your oversight.
Autonomy is on a continuum; you don’t need to sign away all oversight to get a little autonomy.
I don’t know. They don’t need to be inherently honest as long as they follow the instruction: be honest, because that is the first thing I’m telling them every time.
That’s right. Functional instruction-following is surprisingly close to all you need.
The one exception is that you need to somehow get it to prioritize future instructions. An obvious subgoal of fully completing any given instruction is making sure no authorized party gives you a different instruction before you’ve finished.
That is not what we’re training for now. Models currently aren’t smart enough to work through the full logic of every request, but they will be.
You could argue that every instruction has an implicit “unless I change my mind or you misunderstood what I meant,” and to some degree they do. But making sure that subtlety gets through isn’t something we’re training or testing for explicitly yet.
If the model’s goal is to do what you want, and you originally wanted the model to let you change the goals you gave it, then the model ought to understand that (since it’s not dumb) and let you change the goals you gave it..
I think this is intuitive and traditional but wrong.
What people want isn’t autonomy, but do-what-I-mean-and-check (DWIMAC).
Most bosses don’t tell their employees “go do this project” but rather “go plan this project and bounce it off me before you proceed” if the project takes longer than a day or a week. Certainly things like “green the deserts” or “make the world better” are worth a few minutes of your oversight.
Autonomy is on a continuum; you don’t need to sign away all oversight to get a little autonomy.
Notice how doing that well involves things that are suspiciously virtue-like. Most obviously, you’d want the model to be deeply honest.
I don’t know. They don’t need to be inherently honest as long as they follow the instruction: be honest, because that is the first thing I’m telling them every time.
Okay, and how sure are you that the person actually prompting them will remember all the other things like that that are “obvious”?
It is generally implicit in an instruction that you don’t want to simply be mislead about its completion.
That’s right. Functional instruction-following is surprisingly close to all you need.
The one exception is that you need to somehow get it to prioritize future instructions. An obvious subgoal of fully completing any given instruction is making sure no authorized party gives you a different instruction before you’ve finished.
That is not what we’re training for now. Models currently aren’t smart enough to work through the full logic of every request, but they will be.
You could argue that every instruction has an implicit “unless I change my mind or you misunderstood what I meant,” and to some degree they do. But making sure that subtlety gets through isn’t something we’re training or testing for explicitly yet.
If the model’s goal is to do what you want, and you originally wanted the model to let you change the goals you gave it, then the model ought to understand that (since it’s not dumb) and let you change the goals you gave it..
Yes, but I don’t think that’s suspicious. Of course honesty/not reward hacking is a big part of intent alignment.