And even if every time you create a unit of safety research you also create a unit of capabilities research, that’s better than the status quo
There’s a lot of problems with reasoning like this:
No one is certain what chunk of research represents one unit of safety research when in comes to solving the important problems in time, and no one is certain what chunk of research represents one unit of capabilities research when it comes speeding progress to AI catastrophe, especially before doing the research.
If you mean something other than the above by one “unit”, it’s very unclear whether your statement is correct. For example, it’s unclear if it’s net positive to add one additional FTE to each of two teams at Anthropic, one nominally the alignment team and the other nominally the capabilities team. I would argue no, but it’s certainly less obvious than the “one unit for one unit” framing makes it seem.
It’s easy for well meaning people doing work that has dual effects accelerating safety and capabilities to overvalue the impact of the safety work and undervalue the impact of improving capabilities, arising from the general difficulty of fully avoiding motivated reasoning.
It makes it easier for people insufficiently concerned about AI danger, or otherwise having net-negative behaviors or worldviews to blend in with well-meaning people by making some gesture towards “alignment” or “safety”.
Value drift is a real and hard to avoid problem.
The counterfactual to doing dual-purpose alignment/capabilities work is probably not doing/funding nothing helpful, so I don’t know if it’s correct to compare one unit of alignment + capabilities work with the status quo.
In combination with the prior point, there’s a strong social effect, your work is going to strongly influence others’ work, especially when those others are mostly early-mid career.
By being insufficiently cautious about some of the above points, it’s plausibly that we (as in people concerned with x-risk) have collectively squandered some amount of opportunity to build organizations that are robustly net positive, while at the same time greatly speeding capabilities progress.
After writing this out, however, I am wondering to what extent it’s relying on the belief that the last 5-10 years have been generally gone worse than might have been expected when it comes to the trajectory of AI. I think if you’re someone that believes the last 5-10 years have gone about as well as could been expected perhaps its less persuasive.
There’s a lot of problems with reasoning like this:
No one is certain what chunk of research represents one unit of safety research when in comes to solving the important problems in time, and no one is certain what chunk of research represents one unit of capabilities research when it comes speeding progress to AI catastrophe, especially before doing the research.
If you mean something other than the above by one “unit”, it’s very unclear whether your statement is correct. For example, it’s unclear if it’s net positive to add one additional FTE to each of two teams at Anthropic, one nominally the alignment team and the other nominally the capabilities team. I would argue no, but it’s certainly less obvious than the “one unit for one unit” framing makes it seem.
It’s easy for well meaning people doing work that has dual effects accelerating safety and capabilities to overvalue the impact of the safety work and undervalue the impact of improving capabilities, arising from the general difficulty of fully avoiding motivated reasoning.
It makes it easier for people insufficiently concerned about AI danger, or otherwise having net-negative behaviors or worldviews to blend in with well-meaning people by making some gesture towards “alignment” or “safety”.
Value drift is a real and hard to avoid problem.
The counterfactual to doing dual-purpose alignment/capabilities work is probably not doing/funding nothing helpful, so I don’t know if it’s correct to compare one unit of alignment + capabilities work with the status quo.
In combination with the prior point, there’s a strong social effect, your work is going to strongly influence others’ work, especially when those others are mostly early-mid career.
By being insufficiently cautious about some of the above points, it’s plausibly that we (as in people concerned with x-risk) have collectively squandered some amount of opportunity to build organizations that are robustly net positive, while at the same time greatly speeding capabilities progress.
After writing this out, however, I am wondering to what extent it’s relying on the belief that the last 5-10 years have been generally gone worse than might have been expected when it comes to the trajectory of AI. I think if you’re someone that believes the last 5-10 years have gone about as well as could been expected perhaps its less persuasive.