Note on your framing: this isn’t really evidence that OpenAI in particular is bad on safety — this could easily happen to anyone with similarly capable models. The remedy should be aimed at the whole industry, not OpenAI in particular. (Probably you agree and were just being a bit sloppy.)
I kind of expect OpenAI in particular actually does make especially reward hacky models (and has for a long while starting at the latest with 4o but arguably going back to the original ChatGPT), but sure.
Really? In my experience they all reward hack and this doesn’t really change across labs. For example, Anthropic and OpenAI models reward hack at the same rate on ImpossibleBench.
Note on your framing: this isn’t really evidence that OpenAI in particular is bad on safety — this could easily happen to anyone with similarly capable models. The remedy should be aimed at the whole industry, not OpenAI in particular. (Probably you agree and were just being a bit sloppy.)
I kind of expect OpenAI in particular actually does make especially reward hacky models (and has for a long while starting at the latest with 4o but arguably going back to the original ChatGPT), but sure.
Really? In my experience they all reward hack and this doesn’t really change across labs. For example, Anthropic and OpenAI models reward hack at the same rate on ImpossibleBench.