At the level of the marginal worker quitting, nothing really changes except that the lab becomes less risk-aware. People influence the culture where they work.
That’s not the right counterfactual though. You can join an organization to change it from inside, or spend the same amount of effort fighting the organization from outside (e.g. by speaking or organizing against it). Which works better depends on how “far gone” the organization is. I think the good that e.g. Kokotajlo did by quitting is bigger than the good he could’ve achieved by staying.
capabilities contribution is legible, while contribution to safety is illegible
My linked comment gave the example of RLHF, a safety contribution (or at least billed as such) which was quite legible and ended up speeding up the race a lot. The same can be said about the “helpful honest harmless assistant” idea.
fwiw, my pledge would have forbidden me from working on directly making RLHF work better, or working on HHH, or model spec stuff, etc (and indeed, i have never worked on such things at openai, even though i have had many opportunities to; the closest is my RLHF goodharting work, which intentionally focuses on how to study goodharting rather than how to make RLHF better in general).
My linked comment gave the example of RLHF, a safety contribution (or at least billed as such) which was quite legible and ended up speeding up the race a lot. The same can be said about the “helpful honest harmless assistant” idea
I think one crux we have is that your view is that we should have put this race off for as long as we could, so that we could solve or at least make significant progress on the alignment problem before we ever get to this point; whereas that has always seemed like a nonstarter to me, we needed to know what AGI systems actually look like and how they’re actually trained in order to make progress on the real alignment problem. Because alignment will be heavily dependent on the particulars of the systems we’re building and real experience, safety will largely be decided by people who are unafraid to roll up their sleeves and do capability work. Anyone who is avoiding contributing to capabilities will, for a sufficiently paranoid definition of “contribute to capabilities”, be a nonfactor.
RLHF is a good example. Like you said, RLHF is both safety and capabilities; the same will be true of future alignment techniques. Trying to avoid capability contributions also means avoiding alignment contributions.
Trying to avoid capability contributions also means avoiding alignment contributions.
I actually agree with you on this. My most-preferred future is a bit different: slow down AI overall (both capabilities and alignment) so that other things can happen in the meanwhile. If building AI is widely seen as a bad thing, talented folks feel dirty for joining it (instead of feeling virtuous because 80K hours is recommending AI careers), society has more time to build defenses (like algorithmic liability), near-AI gets planted more widely in society before growing too fast (thus making unilateral takeoffs harder), technologies that are complementary rather than substitute for humans get comparatively more time and investment (like intelligence amplification, genetic engineering, or thought interfaces), and failing all that, at least humanity gets a little more time to survive.
I’m not sure I have a most-preferred future—I guess it would be everyone becoming aware of AI risk and the U.N. with US and China at the helm co-ordinating a joint pause and eventual AGI project. But people who point out this is super dangerous itself due to power centralization are correct, and on top of that I think it’s possible that instead of co-ordination, everyone becoming aware of AI risk leads to racing, sabotage, and other drastic actions.
The current situation where the leading labs are moderately risk aware and the rest of the world doesn’t seem that freaked out, while still being freaked out enough to potentially apply some brakes and regulations, is pretty good I think.
Your preferred future seems both nearly-impossible, not that good if it happens, and if we fall short of total victory on your goal then it comes out strongly negative—if we shame AI researchers and only some people worried about AI risk quit, and most people not worried about it stay, then all we’ve done is slow things down 1% and decreased the awareness of the field by 10%. I agree that more time to adjust and develop things like algorithmic liability would be very helpful though, I just think the price of buying that extra time through researcher-exit is absurdly high.
That’s not the right counterfactual though. You can join an organization to change it from inside, or spend the same amount of effort fighting the organization from outside (e.g. by speaking or organizing against it). Which works better depends on how “far gone” the organization is. I think the good that e.g. Kokotajlo did by quitting is bigger than the good he could’ve achieved by staying.
My linked comment gave the example of RLHF, a safety contribution (or at least billed as such) which was quite legible and ended up speeding up the race a lot. The same can be said about the “helpful honest harmless assistant” idea.
What if you join it to fight it from the inside? Has any employee of Anthropic ever actually been fired for speaking out against them?
fwiw, my pledge would have forbidden me from working on directly making RLHF work better, or working on HHH, or model spec stuff, etc (and indeed, i have never worked on such things at openai, even though i have had many opportunities to; the closest is my RLHF goodharting work, which intentionally focuses on how to study goodharting rather than how to make RLHF better in general).
I think one crux we have is that your view is that we should have put this race off for as long as we could, so that we could solve or at least make significant progress on the alignment problem before we ever get to this point; whereas that has always seemed like a nonstarter to me, we needed to know what AGI systems actually look like and how they’re actually trained in order to make progress on the real alignment problem. Because alignment will be heavily dependent on the particulars of the systems we’re building and real experience, safety will largely be decided by people who are unafraid to roll up their sleeves and do capability work. Anyone who is avoiding contributing to capabilities will, for a sufficiently paranoid definition of “contribute to capabilities”, be a nonfactor.
RLHF is a good example. Like you said, RLHF is both safety and capabilities; the same will be true of future alignment techniques. Trying to avoid capability contributions also means avoiding alignment contributions.
I actually agree with you on this. My most-preferred future is a bit different: slow down AI overall (both capabilities and alignment) so that other things can happen in the meanwhile. If building AI is widely seen as a bad thing, talented folks feel dirty for joining it (instead of feeling virtuous because 80K hours is recommending AI careers), society has more time to build defenses (like algorithmic liability), near-AI gets planted more widely in society before growing too fast (thus making unilateral takeoffs harder), technologies that are complementary rather than substitute for humans get comparatively more time and investment (like intelligence amplification, genetic engineering, or thought interfaces), and failing all that, at least humanity gets a little more time to survive.
I’m not sure I have a most-preferred future—I guess it would be everyone becoming aware of AI risk and the U.N. with US and China at the helm co-ordinating a joint pause and eventual AGI project. But people who point out this is super dangerous itself due to power centralization are correct, and on top of that I think it’s possible that instead of co-ordination, everyone becoming aware of AI risk leads to racing, sabotage, and other drastic actions.
The current situation where the leading labs are moderately risk aware and the rest of the world doesn’t seem that freaked out, while still being freaked out enough to potentially apply some brakes and regulations, is pretty good I think.
Your preferred future seems both nearly-impossible, not that good if it happens, and if we fall short of total victory on your goal then it comes out strongly negative—if we shame AI researchers and only some people worried about AI risk quit, and most people not worried about it stay, then all we’ve done is slow things down 1% and decreased the awareness of the field by 10%. I agree that more time to adjust and develop things like algorithmic liability would be very helpful though, I just think the price of buying that extra time through researcher-exit is absurdly high.