I know there’s this kind of evidence, but it’s so discordant with my experience. I use Claude Code/Cowork all day every day for work, and haven’t once it had do anything reward-hackey or strategically deceptive (at least that I can remember? or that I’ve caught?) in the last ~3 months. Do other people actually have the experience of it reward hacking in their day-to-day use?
I’m somewhat suspicious that it might be dependent on usage. I’d be curious if people who more closely supervise its outputs and engage more actively in conversation with it get less reward hacking (since it would know there’s someone home).
I know there’s this kind of evidence, but it’s so discordant with my experience. I use Claude Code/Cowork all day every day for work, and haven’t once it had do anything reward-hackey or strategically deceptive (at least that I can remember? or that I’ve caught?) in the last ~3 months. Do other people actually have the experience of it reward hacking in their day-to-day use?
I’m somewhat suspicious that it might be dependent on usage. I’d be curious if people who more closely supervise its outputs and engage more actively in conversation with it get less reward hacking (since it would know there’s someone home).