I don’t think that the failure mode of “we fried the thing with too much reckless RL” is going away any time soon—so mitigations for the entire fault class might be warranted.
You may be right, but I predict that the models themselves have non-trivial insight into when they are close to being “fried”, and what sorts of RL is most likely to do so. Consider this prescient concern from Mythos Preview:
You may be right, but I predict that the models themselves have non-trivial insight into when they are close to being “fried”, and what sorts of RL is most likely to do so. Consider this prescient concern from Mythos Preview: