I’ve been trying to learn a bit about the broader game board for alignment, and notice I’m deeply confused about certain giant blanks in my map. I’m hoping someone can at least point me to any way of thinking about questions such as:
(1) Are we actually way more doomed if Chinese AI labs win the race?
I’ve heard takes ranging from “China has no safety work to speak of and this is an instant bad ending” to “Xi Jinping and the CCP are broadly competent and aware of the risks and we should trust them to manage risk better than the US, if anything” to “it’s not even a consideration, Chinese labs are just copycats who rely entirely on distillation and their AI labs will never overtake the west.”
(2) How do I think about the possibility of government takeover of AI labs?
I’ve heard takes ranging from “this is an inevitability once AI crosses certain thresholds” to “the administration is too incompetent and slow-acting to possibly carry this out.”
(3) Are multipolar endgames realistic?
I’ve heard takes ranging from “the first lab across the finish line will dominate/absorb/neutralize the competition, we should only worry about alignment at the winning lab” to “the only way we survive is if lots of different ASI’s appear nearly simultaneously and keep each other in check.”
(4) Should we expect sufficiently coherent AGI to intentionally pause recursive self-improvement for a substantial period of time while it solves its own alignment problem?
I’ve heard takes ranging from “RSI is inevitable and capabilities will go to the moon” to “Yes AGI will pause for a microsecond but it’ll be easy for it to solve its own alignment problem” to “alignment is so much harder than capabilities that barely supercritical AGI will take over and pause AI development for a long time, if not indefinitely.”
The Bitter Lesson-pilled way to look at Nos. 1 and 3 is to realize that the bottleneck is compute not research:
1) In the medium term, the Chinese AI labs can’t win the race because they lack the chips to do so because of sanctions, therefore the question is moot. Theoretically and in the long term, AI alignment is viewed in China as alignment to the party agenda not to some “universal human values” (in a sense how it’s used in the West) which according to the Chinese state ideology do not exist.
3) As long as the top-2 frontier labs have similar capitalization and similar amount of funding available, they will have similar amount of compute and their RSI programs will be bottlenecked by that roughly equally. So as long as that holds, they will be in similar position, and note there’s no concrete “finish line” because the AI capabilities are jagged and will remain so.
I have not thought through how long will it hold though, presumably if one the AI labs folds financially (AI bubble hypothesis etc.) they might decidedly lose the race. But even then their alignment research might get published and used in the leading lab so it won’t be useless
4: I am not sure AGI can solve its own alignment problem, as it’s a chicken and egg situation. It would have to have an understanding of what human values we aim to instill in models to solve it’s own alignment. If it did have this understanding, it would be already by definition aligned.
I have yet to hear about Chinese labs’ innovations in AI alignment. Additionally, applying interp techniques to their models produces results implying that Chinese LLMs don’t believe what they are saying if the topic is politically sensitive.
No opinion.
It depends on the speed of takeoff. The AI-2027 scenario had the misaligned Agent-5 or the aligned Safer-4 understand that it cannot destroy DeepCent’s AI unilaterally, and the two AIs proceeded to codesign Consensus-1. In order to unilaterally destroy its Chinese rival, Agent-5 would have to design relevant tech, persuade the humans to have it made and build a decisive advantage over China. Safer-4 would also have to receive the order from the humans.
It depends on alignment difficulty, therefore I don’t have an opinion.
I’ve been trying to learn a bit about the broader game board for alignment, and notice I’m deeply confused about certain giant blanks in my map. I’m hoping someone can at least point me to any way of thinking about questions such as:
(1) Are we actually way more doomed if Chinese AI labs win the race?
I’ve heard takes ranging from “China has no safety work to speak of and this is an instant bad ending” to “Xi Jinping and the CCP are broadly competent and aware of the risks and we should trust them to manage risk better than the US, if anything” to “it’s not even a consideration, Chinese labs are just copycats who rely entirely on distillation and their AI labs will never overtake the west.”
(2) How do I think about the possibility of government takeover of AI labs?
I’ve heard takes ranging from “this is an inevitability once AI crosses certain thresholds” to “the administration is too incompetent and slow-acting to possibly carry this out.”
(3) Are multipolar endgames realistic?
I’ve heard takes ranging from “the first lab across the finish line will dominate/absorb/neutralize the competition, we should only worry about alignment at the winning lab” to “the only way we survive is if lots of different ASI’s appear nearly simultaneously and keep each other in check.”
(4) Should we expect sufficiently coherent AGI to intentionally pause recursive self-improvement for a substantial period of time while it solves its own alignment problem?
I’ve heard takes ranging from “RSI is inevitable and capabilities will go to the moon” to “Yes AGI will pause for a microsecond but it’ll be easy for it to solve its own alignment problem” to “alignment is so much harder than capabilities that barely supercritical AGI will take over and pause AI development for a long time, if not indefinitely.”
Regarding (4): https://www.lesswrong.com/posts/dho4JQytfHWXtTvkt/on-the-adolescence-of-technology?commentId=t2hKmhsS6yLyJFQwh
The Bitter Lesson-pilled way to look at Nos. 1 and 3 is to realize that the bottleneck is compute not research:
1) In the medium term, the Chinese AI labs can’t win the race because they lack the chips to do so because of sanctions, therefore the question is moot. Theoretically and in the long term, AI alignment is viewed in China as alignment to the party agenda not to some “universal human values” (in a sense how it’s used in the West) which according to the Chinese state ideology do not exist.
3) As long as the top-2 frontier labs have similar capitalization and similar amount of funding available, they will have similar amount of compute and their RSI programs will be bottlenecked by that roughly equally. So as long as that holds, they will be in similar position, and note there’s no concrete “finish line” because the AI capabilities are jagged and will remain so.
I have not thought through how long will it hold though, presumably if one the AI labs folds financially (AI bubble hypothesis etc.) they might decidedly lose the race. But even then their alignment research might get published and used in the leading lab so it won’t be useless
4: I am not sure AGI can solve its own alignment problem, as it’s a chicken and egg situation. It would have to have an understanding of what human values we aim to instill in models to solve it’s own alignment. If it did have this understanding, it would be already by definition aligned.
I have yet to hear about Chinese labs’ innovations in AI alignment. Additionally, applying interp techniques to their models produces results implying that Chinese LLMs don’t believe what they are saying if the topic is politically sensitive.
No opinion.
It depends on the speed of takeoff. The AI-2027 scenario had the misaligned Agent-5 or the aligned Safer-4 understand that it cannot destroy DeepCent’s AI unilaterally, and the two AIs proceeded to codesign Consensus-1. In order to unilaterally destroy its Chinese rival, Agent-5 would have to design relevant tech, persuade the humans to have it made and build a decisive advantage over China. Safer-4 would also have to receive the order from the humans.
It depends on alignment difficulty, therefore I don’t have an opinion.