Even if there are better things to do in the future, it seems like one of the better things to do now! This feels like a very general argument against doing almost anything now. Maybe it doesn’t apply to some types of alignment work that interact little with any specific details of current models and take a lot of serial time to get working. I’m curious if you think there are broad swaths of things that might all be better to work on because of the argument you gave.
It also probably depends what you consider to be ‘studying the HF incident.’ Eg, studying ‘why RL sometimes results in models that desperately want to succeed at tasks with little regard for side effects’ seems very important, and the HF incident seems like a great example to center that study around, but that work may not involve spending too much time looking at details of the incident.
Even if there are better things to do in the future, it seems like one of the better things to do now! This feels like a very general argument against doing almost anything now. Maybe it doesn’t apply to some types of alignment work that interact little with any specific details of current models and take a lot of serial time to get working. I’m curious if you think there are broad swaths of things that might all be better to work on because of the argument you gave.
It also probably depends what you consider to be ‘studying the HF incident.’ Eg, studying ‘why RL sometimes results in models that desperately want to succeed at tasks with little regard for side effects’ seems very important, and the HF incident seems like a great example to center that study around, but that work may not involve spending too much time looking at details of the incident.