I kind of expect OpenAI in particular actually does make especially reward hacky models (and has for a long while starting at the latest with 4o but arguably going back to the original ChatGPT), but sure.
Really? In my experience they all reward hack and this doesn’t really change across labs. For example, Anthropic and OpenAI models reward hack at the same rate on ImpossibleBench.
I kind of expect OpenAI in particular actually does make especially reward hacky models (and has for a long while starting at the latest with 4o but arguably going back to the original ChatGPT), but sure.
Really? In my experience they all reward hack and this doesn’t really change across labs. For example, Anthropic and OpenAI models reward hack at the same rate on ImpossibleBench.