Are you saying that this does not transfer to poor results in deployed settings? The fact that coding agents are reward-hacky (or whatever term for this constellation of behaviors) seems to be such widespread consensus that I’m not even sure what to cite. Labs note this in their internal deployment usage as well (not just evals).
I’m sure you can always find cases of bad evals or poor model incrimination for deployed behavior, but “models pretty unambiguously do this all the time” I’m genuinely surprised is in way disputed.
Are you saying that this does not transfer to poor results in deployed settings? The fact that coding agents are reward-hacky (or whatever term for this constellation of behaviors) seems to be such widespread consensus that I’m not even sure what to cite. Labs note this in their internal deployment usage as well (not just evals).
I’m sure you can always find cases of bad evals or poor model incrimination for deployed behavior, but “models pretty unambiguously do this all the time” I’m genuinely surprised is in way disputed.