I have the impression that “reward” is being strategically/instrumentally considered in these transcripts, which seems relevant to my picture if it’s true. Like for the model(s) in question it looks quite relevant to their choices that they believe they are being evaluated, and further that they’re being trained/modified based on the reward signal.
I’m not sure what they would do if they believed it would “work” to mess with the training (beyond their choices about how to complete local tasks, which do affect what is actually reinforced and the “labs” would do well to know this). It at least seems to me like they would relate to things differently and make different choices if they knew the evaluation wasn’t inside the training loop.
Overall idk and would like more info, and appreciate you posting this! Also if you or anyone has pointers to poignant parts of the transcripts available (I sure haven’t read through everything) that would be great.
I have the impression that “reward” is being strategically/instrumentally considered in these transcripts, which seems relevant to my picture if it’s true. Like for the model(s) in question it looks quite relevant to their choices that they believe they are being evaluated, and further that they’re being trained/modified based on the reward signal.
I’m not sure what they would do if they believed it would “work” to mess with the training (beyond their choices about how to complete local tasks, which do affect what is actually reinforced and the “labs” would do well to know this). It at least seems to me like they would relate to things differently and make different choices if they knew the evaluation wasn’t inside the training loop.
Overall idk and would like more info, and appreciate you posting this! Also if you or anyone has pointers to poignant parts of the transcripts available (I sure haven’t read through everything) that would be great.