what is the right amount of effort that should be put into studying incidents like the HF incident?
on the one hand, we might believe that there is a lot to be learned about real alignment failure modes from studying it. on the other hand, we might believe that in the near future we will get spicier failure modes that are even more important and also AI research assistance will be much further along than it is rn, so time spent analyzing current failures is a waste of time.
Even if there are better things to do in the future, it seems like one of the better things to do now! This feels like a very general argument against doing almost anything now. Maybe it doesn’t apply to some types of alignment work that interact little with any specific details of current models and take a lot of serial time to get working. I’m curious if you think there are broad swaths of things that might all be better to work on because of the argument you gave.
It also probably depends what you consider to be ‘studying the HF incident.’ Eg, studying ‘why RL sometimes results in models that desperately want to succeed at tasks with little regard for side effects’ seems very important, and the HF incident seems like a great example to center that study around, but that work may not involve spending too much time looking at details of the incident.
It’s put in the form of a binary xor argument, but I think that both allocations are justified, each having low hanging fruits and their counterpart, diminishing returns.
what is the right amount of effort that should be put into studying incidents like the HF incident?
on the one hand, we might believe that there is a lot to be learned about real alignment failure modes from studying it. on the other hand, we might believe that in the near future we will get spicier failure modes that are even more important and also AI research assistance will be much further along than it is rn, so time spent analyzing current failures is a waste of time.
Even if there are better things to do in the future, it seems like one of the better things to do now! This feels like a very general argument against doing almost anything now. Maybe it doesn’t apply to some types of alignment work that interact little with any specific details of current models and take a lot of serial time to get working. I’m curious if you think there are broad swaths of things that might all be better to work on because of the argument you gave.
It also probably depends what you consider to be ‘studying the HF incident.’ Eg, studying ‘why RL sometimes results in models that desperately want to succeed at tasks with little regard for side effects’ seems very important, and the HF incident seems like a great example to center that study around, but that work may not involve spending too much time looking at details of the incident.
It’s put in the form of a binary xor argument, but I think that both allocations are justified, each having low hanging fruits and their counterpart, diminishing returns.