Useful distinction, and I agree with a lot of this. A few complications, though.
As others have said here, better understanding of AI behavior/cognition will also help with current alignment goals, so the science speeds things up too. Knowing why a method works is the best route to a better method.
I think it’s very hard to tell which techniques will still help when we hit weak superintelligence, let alone strong. Current alignment techniques may very well keep helping if LLMs get us at least to weak superintelligence. I doubt they’re sufficient, but they may help. So I don’t think the engineering vs. science line tells us which work will turn out to matter.
We are probably going to rely heavily on AI to help with alignment. We’ll try to use it for conceptual questions whether or not the answers are likely to be slop. If we could reduce slop without aiding general reasoning, it would pretty clearly be beneficial. Better metacognition is how humans reduce our slop, but it also helps our capabilities a lot, because knowing when you’re not sure lets you work harder on that part of the problem.
I don’t think there’s a way to avoid dilemmas like this. Capabilities are going to keep improving with or without the help of the risk-concerned. Judging whether we’re helping the odds of alignment more than we’re hurting with acceleration is going to remain hard. Each case should probably be publicly discussed in detail. It’s easy to let motivated reasoning convince you that your favorite alignment approach will be more good than harm. But doing nothing to keep our hands clean is not likely to be the best move, either.
Useful distinction, and I agree with a lot of this. A few complications, though.
As others have said here, better understanding of AI behavior/cognition will also help with current alignment goals, so the science speeds things up too. Knowing why a method works is the best route to a better method.
I think it’s very hard to tell which techniques will still help when we hit weak superintelligence, let alone strong. Current alignment techniques may very well keep helping if LLMs get us at least to weak superintelligence. I doubt they’re sufficient, but they may help. So I don’t think the engineering vs. science line tells us which work will turn out to matter.
The capability/alignment dilemma I’ve been most concerned with is on a different axis. I think that Human-like metacognitive skills will reduce LLM slop and aid alignment and capabilities. I’ve been wondering whether the capabilities gain might be worth it.
We are probably going to rely heavily on AI to help with alignment. We’ll try to use it for conceptual questions whether or not the answers are likely to be slop. If we could reduce slop without aiding general reasoning, it would pretty clearly be beneficial. Better metacognition is how humans reduce our slop, but it also helps our capabilities a lot, because knowing when you’re not sure lets you work harder on that part of the problem.
I don’t think there’s a way to avoid dilemmas like this. Capabilities are going to keep improving with or without the help of the risk-concerned. Judging whether we’re helping the odds of alignment more than we’re hurting with acceleration is going to remain hard. Each case should probably be publicly discussed in detail. It’s easy to let motivated reasoning convince you that your favorite alignment approach will be more good than harm. But doing nothing to keep our hands clean is not likely to be the best move, either.