To the extent that AGI alignment is difficult[1], I think “misalignment science” is a poor term for a field of inquiry that aims to get at the meat of what technical alignment should aim to get at. This is like how it’d be silly to frame the study of how to make lasers as “failure-for-a-random-clump-of-stuff-to-be-a-laser science”. Like, if we want to make lasers, we should ask for [fundamental understanding of light and of materials and understanding of certain specific phenomena and design and fabrication ideas and methods and protocols and specifications of component designs etc] relevant to making lasers — that is, for fields of optics and quantum theory and condensed matter physics and laser science and engineering etc — whereas it makes much less sense to ask for [an understanding of the myriad ways in which a collection of fundamental particles might fail to make up a laser]. Ditto for lenses and particle accelerators and microchips and technologies generally.
And I do think AGI alignment is probably extremely difficult, and so by modus ponens, I think “misalignment science” is a poor term. But I sorta don’t want to defend this part in the present comment. I mostly just want to point out that [using the term “misalignment science” to designate a field of inquiry that captures a lot of what alignment ought to be doing] would contribute to framing the AGI situation as one where there is “normal good/benign behavior” and one studies pathological deviations away from it (for example, medicine has historically been such a field[2]).
That said, I think that your three-point characterization of “misalignment science” is not actually making this imo-mistake. I agree understanding is good and important in AI alignment (though purely technically, ie ignoring governance implications, I’m also very pro alignment engineering work[3]). I agree that security mindset is good and important in AI alignment, and that people would do well to think more about ways their plans might fail[4]. I agree alignment research success should not need to look like having an immediate very practical use case.
But this doesn’t make “misalignment science” a fine term. Furthermore, looking at the sub-approaches you list after your characterization, I think you probably are also making an error in content in this direction indicated by the imo-error in your choice of term.[5]
ie to the extent that it is difficult to have a good transition to a world with artificial generally intelligent systems, or that it is difficult to make AI systems that it is good/fine to hand the future to, or that it is difficult to make AI systems that extend a human’s agency or just benignly do tasks or provide services to humanity, or whatever
I’d only give sth like that the main reason we won’t be disempowered is the present stack of fairly shallow ML research (though tbh I’d include basically everything you list under “misalignment science” under “shallow ML research” as well, but I’d consider alignment engineering to get most of the Shapley here — anyway there are probably disagreements here that are outside the intended scope of the present comment), but I’d also only give sth like that we won’t be disempowered mostly due to other things. (In other words, I’m saying: like we will be disempowered, but ML slop gets credit for a bunch of the remaining .)
Btw I feel like your explicit three-point characterization of “misalignment science” and your list of sub-approaches+examples are quite incongruent, at least if I am to take the latter as aiming to span the space decently well. This makes me mostly suspect that your explicit characterization isn’t getting at the field you really have in mind.
To the extent that AGI alignment is difficult[1], I think “misalignment science” is a poor term for a field of inquiry that aims to get at the meat of what technical alignment should aim to get at. This is like how it’d be silly to frame the study of how to make lasers as “failure-for-a-random-clump-of-stuff-to-be-a-laser science”. Like, if we want to make lasers, we should ask for [fundamental understanding of light and of materials and understanding of certain specific phenomena and design and fabrication ideas and methods and protocols and specifications of component designs etc] relevant to making lasers — that is, for fields of optics and quantum theory and condensed matter physics and laser science and engineering etc — whereas it makes much less sense to ask for [an understanding of the myriad ways in which a collection of fundamental particles might fail to make up a laser]. Ditto for lenses and particle accelerators and microchips and technologies generally.
And I do think AGI alignment is probably extremely difficult, and so by modus ponens, I think “misalignment science” is a poor term. But I sorta don’t want to defend this part in the present comment. I mostly just want to point out that [using the term “misalignment science” to designate a field of inquiry that captures a lot of what alignment ought to be doing] would contribute to framing the AGI situation as one where there is “normal good/benign behavior” and one studies pathological deviations away from it (for example, medicine has historically been such a field[2]).
That said, I think that your three-point characterization of “misalignment science” is not actually making this imo-mistake. I agree understanding is good and important in AI alignment (though purely technically, ie ignoring governance implications, I’m also very pro alignment engineering work[3]). I agree that security mindset is good and important in AI alignment, and that people would do well to think more about ways their plans might fail[4]. I agree alignment research success should not need to look like having an immediate very practical use case.
But this doesn’t make “misalignment science” a fine term. Furthermore, looking at the sub-approaches you list after your characterization, I think you probably are also making an error in content in this direction indicated by the imo-error in your choice of term.[5]
ie to the extent that it is difficult to have a good transition to a world with artificial generally intelligent systems, or that it is difficult to make AI systems that it is good/fine to hand the future to, or that it is difficult to make AI systems that extend a human’s agency or just benignly do tasks or provide services to humanity, or whatever
I say “historically” because I expect this would change fairly soon in a humane future, with people getting more into longevity/”healthmaxxing”.
I’d only give sth like that the main reason we won’t be disempowered is the present stack of fairly shallow ML research (though tbh I’d include basically everything you list under “misalignment science” under “shallow ML research” as well, but I’d consider alignment engineering to get most of the Shapley here — anyway there are probably disagreements here that are outside the intended scope of the present comment), but I’d also only give sth like that we won’t be disempowered mostly due to other things. (In other words, I’m saying: like we will be disempowered, but ML slop gets credit for a bunch of the remaining .)
Though I want to note that in practice, in the current AI safety community, this will largely look like looking for causes of misalignment exclusively or even just mainly in training pressures, which imo entails already having fundamentally misunderstood the nature of AGI risk. I believe this is a long-standing disagreement that I won’t do justice to in the present comment, but I say some relevant things in this comment.
Btw I feel like your explicit three-point characterization of “misalignment science” and your list of sub-approaches+examples are quite incongruent, at least if I am to take the latter as aiming to span the space decently well. This makes me mostly suspect that your explicit characterization isn’t getting at the field you really have in mind.