Fwiw I basically agree with this. When I came to write up the list I wanted something concrete I could point to, but struggled to find any really good examples. I don’t think this is clear from the text.
In general, when you’re advocating something and you can’t find good examples of it, that should make you question whether that’s the right thing to advocate for at all.
I do take your point that you were distinguishing misalignment science from ML more than I gave you credit for; sorry about that. I think you should go further with this, and characterize what you wanted using examples of the best science you know from outside alignment, to point towards the thing you’re excited about. Doing this might have led you to rename “misalignment science”, because misalignment isn’t fundamental enough to be the main focus of a science (it feels like calling neuroscience “brain disorder science”, or chemistry “explosion science”).
I would defend the works under “Pushing understanding” as directionally better though, which was my main aim.
The problem with “directionally correct” is that it erodes our ability to draw category boundaries. For example, I want people building cool products more than I want people scaling up neural networks. But I shouldn’t call the former “alignment research”, even though it’s directionally good for people to shift that way.
Lately I’ve been trying to shift people from doing pragmatic AI alignment research to being much more scientific. But my sense is that most people doing such research could easily relabel themselves as doing “misalignment science” as you’ve described it, while changing their research relatively little (e.g. doing the same thing but adding more post-hoc analysis). Hence it erodes the thing I’m trying to gesture at with the word “science” (despite a bunch of your other arguments making good and important points).
So if I can convince a bunch of “number go up” people to instead do more scientific work, my guess is that the second order effect is that more people also end up doing conceptual / theoretical work.
This comment you made below feels like a crux to me, because of the eroding categories thing I talked about above.
In general, when you’re advocating something and you can’t find good examples of it, that should make you question whether that’s the right thing to advocate for at all.
I do take your point that you were distinguishing misalignment science from ML more than I gave you credit for; sorry about that. I think you should go further with this, and characterize what you wanted using examples of the best science you know from outside alignment, to point towards the thing you’re excited about. Doing this might have led you to rename “misalignment science”, because misalignment isn’t fundamental enough to be the main focus of a science (it feels like calling neuroscience “brain disorder science”, or chemistry “explosion science”).
The problem with “directionally correct” is that it erodes our ability to draw category boundaries. For example, I want people building cool products more than I want people scaling up neural networks. But I shouldn’t call the former “alignment research”, even though it’s directionally good for people to shift that way.
Lately I’ve been trying to shift people from doing pragmatic AI alignment research to being much more scientific. But my sense is that most people doing such research could easily relabel themselves as doing “misalignment science” as you’ve described it, while changing their research relatively little (e.g. doing the same thing but adding more post-hoc analysis). Hence it erodes the thing I’m trying to gesture at with the word “science” (despite a bunch of your other arguments making good and important points).
This comment you made below feels like a crux to me, because of the eroding categories thing I talked about above.