So, are they going to take all those examples and send through ‘eyyy this was bad instead of good’ gradients during a quick fine tuning session? I guess that’d require re-running the samples and hoping the model cheated again so you could give it negative reward. They need to get better at model monitoring so the labels they give during training are better.
So, are they going to take all those examples and send through ‘eyyy this was bad instead of good’ gradients during a quick fine tuning session? I guess that’d require re-running the samples and hoping the model cheated again so you could give it negative reward. They need to get better at model monitoring so the labels they give during training are better.