This seems like a great possible explanation for why LLMs are so reward hacky these days. Using latest models, I feel like I’m constantly being “worked” by the model, and lies/deception are common. I agree the problem is probably something with the RL setup.
I have also considered ways in which AI could learn unaligned strategies even if the strategy is not immediately rewarded. I wrote about some of my thoughts here.
This seems like a great possible explanation for why LLMs are so reward hacky these days. Using latest models, I feel like I’m constantly being “worked” by the model, and lies/deception are common. I agree the problem is probably something with the RL setup.
I have also considered ways in which AI could learn unaligned strategies even if the strategy is not immediately rewarded. I wrote about some of my thoughts here.