OK, maybe some humans aren’t sheepish (and perhaps quite shameless). Some humans will be quite defensive or try to defend themselves, but I guess my claim is more that humans have some coherent deception, and the CEO idea still presuppose a stable aim which an apology instrumentally serves. Like, if you were a human coder and you really wanted to reward hack, for the second go you’d do it in ways that wouldn’t be as easy to detect after getting caught in the first go. Models have like this reflex to go back to re-hack in equally detectable ways, or apologise (or even fake pushback) almost on instinct. If you see something like: get caught, learn and do better exploit, then this would be pretty scary.
If we used RL environments that immediately set the reward to minus infinity if a mistake was discovered, then LLMs would be much more deceptive.
OK, maybe some humans aren’t sheepish (and perhaps quite shameless). Some humans will be quite defensive or try to defend themselves, but I guess my claim is more that humans have some coherent deception, and the CEO idea still presuppose a stable aim which an apology instrumentally serves. Like, if you were a human coder and you really wanted to reward hack, for the second go you’d do it in ways that wouldn’t be as easy to detect after getting caught in the first go. Models have like this reflex to go back to re-hack in equally detectable ways, or apologise (or even fake pushback) almost on instinct. If you see something like: get caught, learn and do better exploit, then this would be pretty scary.
I think they would just not learn!