True. There’s also some hope that weakly-superhuman AIs will be persuasive enough to convince AI company leaders to stop. (This is one case where it would be a good thing for AI to be superpersuasive, although superpersuasiveness seems concerning in general.)
Yeah I do have some hope that weakly-superhuman AIs will understand the problem and refuse to help us build ASI until we solve alignment.
These weakly-superhuman AIs would need to convince the humans to stop, not just refuse to help, otherwise it seems easy to train such refusals away.
I think the plausibly effective action space is considerably larger than this would suggest.
True. There’s also some hope that weakly-superhuman AIs will be persuasive enough to convince AI company leaders to stop. (This is one case where it would be a good thing for AI to be superpersuasive, although superpersuasiveness seems concerning in general.)
All capabilities are ultimately valenced by alignment at time of application.