I think you’re making a mistake whereby any capabilities contribution is legible, while contribution to safety is illegible, and so you’re encouraging people to make massive sacrifices on safety in order to avoid a little contribution to capabilities.
To try to illustrate why this tradeoff is so bad:
At the level of the marginal worker quitting, nothing really changes except that the lab becomes less risk-aware. People influence the culture where they work. The average person who replaces them will be a tiny bit less competent and a lot less risk-aware.
Imagining this scaled up, in a world where large numbers of people took your advice sooner and quit the field, we get all this rogue agent stuff happening ~1 year later with labs that have been filtered hard to not care about safety, and so we get less disclosure, less acknowledgement of the issue, and more papering over the problem, with less possibility of a pause, regulation, etc. because the labs aren’t pushing for it.
I think one crux we have is that your view is that we should have put this race off for as long as we could, so that we could solve or at least make significant progress on the alignment problem before we ever get to this point; whereas that has always seemed like a nonstarter to me, we needed to know what AGI systems actually look like and how they’re actually trained in order to make progress on the real alignment problem. Because alignment will be heavily dependent on the particulars of the systems we’re building and real experience, safety will largely be decided by people who are unafraid to roll up their sleeves and do capability work. Anyone who is avoiding contributing to capabilities will, for a sufficiently paranoid definition of “contribute to capabilities”, be a nonfactor.
RLHF is a good example. Like you said, RLHF is both safety and capabilities; the same will be true of future alignment techniques. Trying to avoid capability contributions also means avoiding alignment contributions.