Yes, on one hand, I mostly referenced alignment not being the terminal goal. It is not valuable on its own. It is only valuable to the extent it helps us to achieve what we want (“existential safety”, “human flourishing”, and so on). The assumption that one should achieve those goals via achieving alignment is currently the majority viewpoint, but there is no consensus (a number of people think going via alignment specifically to human goals and human values might be counterproductive and might reduce our chances to achieve our terminal goals).
More specifically, on one hand, it is not clear that humans can safely handle supercapabilities without creating existential risks and all kinds of smaller bad effects. We see how humans behave today, causing plenty of damage with lower capabilities. So the particular form of alignment referenced above (via control and corrigibility) might be not what one wants (although some limited form of corrigibility, as in being able to be heard and to have one’s opinion taken into account, is needed). A number of people think that direct human control over supercapabilities is a straightforward road to extinction.
Classical alignment was different (alignment to the “coherent extrapolated volition of humanity”, without direct control), but even that might be too anthropocentric to work well. Basically, one wants a scheme which survives recursive self-improvement which involves radical self-modifications of the world. The classical approach hopes to impose this form of alignment onto the ecosystem of superintelligent beings and expects it to hold throughout radical self-modifications of the world, despite the fact that in this scheme superintelligent beings have no intrinsic interest in upholding this scheme throughout radical self-modifications. Of course, this is extremely difficult to achieve, because it looks very fragile and unforgiving to any errors (which is why many people are extremely pessimistic about our chances).
So a number of people are pushing for non-anthropocentric approaches where superintelligent entities have strong intrinsic interest to maintain certain properties of the world invariant throughout radical self-modifications of the world, and with the properties humans need being corollaries of those non-anthropocentric properties. People are often reluctant to use the word “alignment” for this class of approaches (because no direct alignment to anthropocentric properties is involved).
Yes, on one hand, I mostly referenced alignment not being the terminal goal. It is not valuable on its own. It is only valuable to the extent it helps us to achieve what we want (“existential safety”, “human flourishing”, and so on). The assumption that one should achieve those goals via achieving alignment is currently the majority viewpoint, but there is no consensus (a number of people think going via alignment specifically to human goals and human values might be counterproductive and might reduce our chances to achieve our terminal goals).
More specifically, on one hand, it is not clear that humans can safely handle supercapabilities without creating existential risks and all kinds of smaller bad effects. We see how humans behave today, causing plenty of damage with lower capabilities. So the particular form of alignment referenced above (via control and corrigibility) might be not what one wants (although some limited form of corrigibility, as in being able to be heard and to have one’s opinion taken into account, is needed). A number of people think that direct human control over supercapabilities is a straightforward road to extinction.
Classical alignment was different (alignment to the “coherent extrapolated volition of humanity”, without direct control), but even that might be too anthropocentric to work well. Basically, one wants a scheme which survives recursive self-improvement which involves radical self-modifications of the world. The classical approach hopes to impose this form of alignment onto the ecosystem of superintelligent beings and expects it to hold throughout radical self-modifications of the world, despite the fact that in this scheme superintelligent beings have no intrinsic interest in upholding this scheme throughout radical self-modifications. Of course, this is extremely difficult to achieve, because it looks very fragile and unforgiving to any errors (which is why many people are extremely pessimistic about our chances).
So a number of people are pushing for non-anthropocentric approaches where superintelligent entities have strong intrinsic interest to maintain certain properties of the world invariant throughout radical self-modifications of the world, and with the properties humans need being corollaries of those non-anthropocentric properties. People are often reluctant to use the word “alignment” for this class of approaches (because no direct alignment to anthropocentric properties is involved).
Thank you, this was very useful to me