I’m reminded of this quick take by Daniel Kokotajlo a year ago:
I used to think reward was not going to be the optimization target. I remember hearing Paul Christiano say something like “The AGIs, they are going to crave reward. Crave it so badly,” and disagreeing.
The situationally aware reward hacking results of the past half-year are making me update more towards Paul’s position. Maybe reward (i.e. reinforcement) will increasingly become the optimization target, as RL on LLMs is scaled up massively. Maybe the models will crave reward.
What are the implications of this, if true?
Well, we could end up in Control World: A world where it’s generally understood across the industry that the AIs are not, in fact, aligned, and that they will totally murder you if they think that doing so would get them reinforced. Companies will presumably keep barrelling forward regardless, making their AIs smarter and smarter and having them do more and more coding etc.… but they might put lots of emphasis on having really secure sandboxes for the AIs to operate in, with really hard-to-hack evaluation metrics, possibly even during deployment. “The AI does not love us, but we have a firm grip on its food supply” basically.
Or maybe not; maybe confusion would reign and people would continue to think that the models are aligned and e.g. wouldn’t hurt a fly in real life, they only do it in tests because they know it’s a test.
Or maybe we’d briefly be in Control World until, motivated by economic pressure, the companies come up with some fancier training scheme or architecture that stops the models from learning to crave reinforcement. I wonder what that would be.
I’m reminded of this quick take by Daniel Kokotajlo a year ago: