This looks somewhat similar to the persona adoption or reward hacking (a user wants me to agree with them, so I’ll just go along with everything they say) but I’m unsure
But another question is whether a model playing along the user is metagaming? For example, a model might just have learned to agree with users during RLHF, and it’s not gaming anything. It just learned what users want and plays along, and the fact that some users dislike that doesn’t really make it metagaming, many users like it. But you can also argue that it exploits human biases and reward hacks to them.
Honestly, I’m unsure about all the statements in this comment
Right. It does seem fair to think of this of adopting a particular persona, where the persona is summoned the subtle context. However, this specific mechanism of summoning seems worth singling out.
(FWIW, I was imagining scenarios where this does result in what would reasonably be called metagaming. I guess it wouldn’t show up always—mostly during scheming evaluations, and maybe sometimes during actual use by people who expect the AI to scheme.)
This looks somewhat similar to the persona adoption or reward hacking (a user wants me to agree with them, so I’ll just go along with everything they say) but I’m unsure
But another question is whether a model playing along the user is metagaming? For example, a model might just have learned to agree with users during RLHF, and it’s not gaming anything. It just learned what users want and plays along, and the fact that some users dislike that doesn’t really make it metagaming, many users like it. But you can also argue that it exploits human biases and reward hacks to them.
Honestly, I’m unsure about all the statements in this comment
Right. It does seem fair to think of this of adopting a particular persona, where the persona is summoned the subtle context. However, this specific mechanism of summoning seems worth singling out. (FWIW, I was imagining scenarios where this does result in what would reasonably be called metagaming. I guess it wouldn’t show up always—mostly during scheming evaluations, and maybe sometimes during actual use by people who expect the AI to scheme.)