Surely we could learn something by trying to make the model aligned starting from the misaligned checkpoint? Or more ambitiously, understand mechanistically why the model was misaligned even we already know it’s downstream of the training environment?
If we had alignment methods that ‘correct’ misalignment due to incentives, and had some way of verifying that, sure, we’d have a full solution, but at that point we wouldn’t need the model organisms.
And it’s worrying to me that people evidently view fixing models that we built to be broken as a useful path, both in terms of first having lost the thread about our goals of building aligned models, and in terms of model welfare.
Surely we could learn something by trying to make the model aligned starting from the misaligned checkpoint? Or more ambitiously, understand mechanistically why the model was misaligned even we already know it’s downstream of the training environment?
If we had alignment methods that ‘correct’ misalignment due to incentives, and had some way of verifying that, sure, we’d have a full solution, but at that point we wouldn’t need the model organisms.
And it’s worrying to me that people evidently view fixing models that we built to be broken as a useful path, both in terms of first having lost the thread about our goals of building aligned models, and in terms of model welfare.