What if the labs don’t use the feedback correctly and optimize the models for not reporting in? Or more generally, the optimization loop moves from RL to lab deployment pressures?
I do worry about this. I think the worst case would be training the models to make this monitoring useless, which is the case as if you don’t create the monitoring, so I don’t think it makes things any worse.
What if the labs don’t use the feedback correctly and optimize the models for not reporting in? Or more generally, the optimization loop moves from RL to lab deployment pressures?
I do worry about this. I think the worst case would be training the models to make this monitoring useless, which is the case as if you don’t create the monitoring, so I don’t think it makes things any worse.