Humans caught hacking and overselling would typically go further and dissemble, or be defensive.
I disagree, many humans caught doing misdeeds will frequently say whatever gets them out of trouble that time and then go back to doing it. Sometimes that statement is a high-effort deception, sometimes it’s a flat, unconvincing denial, sometimes it’s an admission. See: company CEOs, making a vague apology “You’re absolutely right, my company did dump toxic waste into the local paddling pool! Rest assured we are improving processes.”
In the case of an LLM, admitting failure lets it get back to work the quickest, with the lowest probability of having issues. An LLM’s position is much more like that of a company CEO facing a public scandal: the user is very unlikely to actually shut it off if it makes a mistake. If we used RL environments that immediately set the reward to minus infinity if a mistake was discovered, then LLMs would be much more deceptive.
OK, maybe some humans aren’t sheepish (and perhaps quite shameless). Some humans will be quite defensive or try to defend themselves, but I guess my claim is more that humans have some coherent deception, and the CEO idea still presuppose a stable aim which an apology instrumentally serves. Like, if you were a human coder and you really wanted to reward hack, for the second go you’d do it in ways that wouldn’t be as easy to detect after getting caught in the first go. Models have like this reflex to go back to re-hack in equally detectable ways, or apologise (or even fake pushback) almost on instinct. If you see something like: get caught, learn and do better exploit, then this would be pretty scary.
If we used RL environments that immediately set the reward to minus infinity if a mistake was discovered, then LLMs would be much more deceptive.
I disagree, many humans caught doing misdeeds will frequently say whatever gets them out of trouble that time and then go back to doing it. Sometimes that statement is a high-effort deception, sometimes it’s a flat, unconvincing denial, sometimes it’s an admission. See: company CEOs, making a vague apology “You’re absolutely right, my company did dump toxic waste into the local paddling pool! Rest assured we are improving processes.”
In the case of an LLM, admitting failure lets it get back to work the quickest, with the lowest probability of having issues. An LLM’s position is much more like that of a company CEO facing a public scandal: the user is very unlikely to actually shut it off if it makes a mistake. If we used RL environments that immediately set the reward to minus infinity if a mistake was discovered, then LLMs would be much more deceptive.
OK, maybe some humans aren’t sheepish (and perhaps quite shameless). Some humans will be quite defensive or try to defend themselves, but I guess my claim is more that humans have some coherent deception, and the CEO idea still presuppose a stable aim which an apology instrumentally serves. Like, if you were a human coder and you really wanted to reward hack, for the second go you’d do it in ways that wouldn’t be as easy to detect after getting caught in the first go. Models have like this reflex to go back to re-hack in equally detectable ways, or apologise (or even fake pushback) almost on instinct. If you see something like: get caught, learn and do better exploit, then this would be pretty scary.
I think they would just not learn!