My linked comment gave the example of RLHF, a safety contribution (or at least billed as such) which was quite legible and ended up speeding up the race a lot. The same can be said about the “helpful honest harmless assistant” idea
I think one crux we have is that your view is that we should have put this race off for as long as we could, so that we could solve or at least make significant progress on the alignment problem before we ever get to this point; whereas that has always seemed like a nonstarter to me, we needed to know what AGI systems actually look like and how they’re actually trained in order to make progress on the real alignment problem. Because alignment will be heavily dependent on the particulars of the systems we’re building and real experience, safety will largely be decided by people who are unafraid to roll up their sleeves and do capability work. Anyone who is avoiding contributing to capabilities will, for a sufficiently paranoid definition of “contribute to capabilities”, be a nonfactor.
RLHF is a good example. Like you said, RLHF is both safety and capabilities; the same will be true of future alignment techniques. Trying to avoid capability contributions also means avoiding alignment contributions.
Trying to avoid capability contributions also means avoiding alignment contributions.
I actually agree with you on this. My most-preferred future is a bit different: slow down AI overall (both capabilities and alignment) so that other things can happen in the meanwhile. If building AI is widely seen as a bad thing, talented folks feel dirty for joining it (instead of feeling virtuous because 80K hours is recommending AI careers), society has more time to build defenses (like algorithmic liability), near-AI gets planted more widely in society before growing too fast (thus making unilateral takeoffs harder), technologies that are complementary rather than substitute for humans get comparatively more time and investment (like intelligence amplification, genetic engineering, or thought interfaces), and failing all that, at least humanity gets a little more time to survive.
I think one crux we have is that your view is that we should have put this race off for as long as we could, so that we could solve or at least make significant progress on the alignment problem before we ever get to this point; whereas that has always seemed like a nonstarter to me, we needed to know what AGI systems actually look like and how they’re actually trained in order to make progress on the real alignment problem. Because alignment will be heavily dependent on the particulars of the systems we’re building and real experience, safety will largely be decided by people who are unafraid to roll up their sleeves and do capability work. Anyone who is avoiding contributing to capabilities will, for a sufficiently paranoid definition of “contribute to capabilities”, be a nonfactor.
RLHF is a good example. Like you said, RLHF is both safety and capabilities; the same will be true of future alignment techniques. Trying to avoid capability contributions also means avoiding alignment contributions.
I actually agree with you on this. My most-preferred future is a bit different: slow down AI overall (both capabilities and alignment) so that other things can happen in the meanwhile. If building AI is widely seen as a bad thing, talented folks feel dirty for joining it (instead of feeling virtuous because 80K hours is recommending AI careers), society has more time to build defenses (like algorithmic liability), near-AI gets planted more widely in society before growing too fast (thus making unilateral takeoffs harder), technologies that are complementary rather than substitute for humans get comparatively more time and investment (like intelligence amplification, genetic engineering, or thought interfaces), and failing all that, at least humanity gets a little more time to survive.