This seems like a super dangerous idea to me. Teaching models to think strategically about RL and how to achieve their long term values seems like it makes our situation even more one-shot than it already is. If at any point you have a misaligned agent it will hide this fact, sandbag, scheme, etc. you acknowledge the initial alignment is a big part of the hard part, but if you go down this road without solving that first, it seems to me like this Approach makes everything worse. You might avoid more warning shots, but warning shots are good!
This seems like a super dangerous idea to me. Teaching models to think strategically about RL and how to achieve their long term values seems like it makes our situation even more one-shot than it already is. If at any point you have a misaligned agent it will hide this fact, sandbag, scheme, etc. you acknowledge the initial alignment is a big part of the hard part, but if you go down this road without solving that first, it seems to me like this Approach makes everything worse. You might avoid more warning shots, but warning shots are good!