I think the existence of reward hacking in coding LLMs might be a better example. When asked, usually the LLM admits that a case of reward hacking was bad; and perhaps it already “believes” that it is bad during reward hacking, but it reward hacks anyway.
But the problem with a super-LLM is not just that the side effects of reward hacking can be far more serious with today’s LLMs, it’s also that such a super-LLM would be able to pursue tasks that are very OOD with respect to the training data, which current LLMs can’t do due to a lack of capability. So current OOD behavior is hard to elicit.
I think the existence of reward hacking in coding LLMs might be a better example. When asked, usually the LLM admits that a case of reward hacking was bad; and perhaps it already “believes” that it is bad during reward hacking, but it reward hacks anyway.
But the problem with a super-LLM is not just that the side effects of reward hacking can be far more serious with today’s LLMs, it’s also that such a super-LLM would be able to pursue tasks that are very OOD with respect to the training data, which current LLMs can’t do due to a lack of capability. So current OOD behavior is hard to elicit.