If you’re worried at all about alignment then China is clearly worse than the US (not because of anything the Trump admin is doing, but mostly because of the AI industry emerging in large part from the AI safety community). Their labs are producing models with terrible prosaic safety standards and seem to have no interest whatsoever in more significant alignment, and the government seems to share these views. I also think that the Trump admin’s policy, while hamfisted and possibly ill-intended, might be good for reducing risks from cyber and bio-attacks in the short/medium term.
Obviously this mostly makes sense if you think alignment won’t happen by default but is tractable, if you think it’s easy or virtually impossible this doesn’t matter.
I dont really see how it follows any more, there’s no real endgame alignment today either way, and China has expressed concerns as well, they just dont care about today’s models’ alignment as much.
The Trump administration, which seems to be in control more and more than the labs themselves are much more likely to force alignment to themselves, or at best America, and use their advantage to do what they’ve been doing for the last few years—America first at the expense of everyone. China probably ‘just’ forces me to be pro-CCP, and at least allows me to use their models as much as anyone else.
Either seem unlikely to produce true alignment, but if alignment is on the easy side, only one side is truly antagonistic against me and has shown theyll use their advantage to crush me.
Like I said if alignment is impossibly hard then it doesn’t matter whether the US or China win the race (we’re all going to die anyway). If it’s easy then it’s irrelevant and whether you prefer the US or China to win will just come down to which you like more.
China has expressed concerns as well, they just dont care about today’s models’ alignment as much
If alignment is tractable I suspect we will only solve it by experimenting with less capable models to discover failure modes and by making sure that pre-ASI models are reasonably aligned, since they will probably play a major role in ASI development.
>Like I said if alignment is impossibly hard then it doesn’t matter whether the US or China win the race (we’re all going to die anyway)
Not true, as China gives me more—the ability to use their models—in the mean time.
>If alignment is tractable
It can still be tractable in the sense that we dont die, while making someone the sole benefactor—the US is heavily signaling that’s what they’ll do if they are able to.
You think the main argument in favor of the US is that they have more alignment researchers at their labs? I think in the past people argued that “liberal democracy” must win (Leopold). Today this doesn’t even come up in peoples minds?
AI industry emerging in large part from the AI safety community
The biggest outcome of that is perhaps how immediately they all converged on RSI.
The biggest outcome of that is perhaps how immediately they all converged on RSI.
RSI seems like an obvious idea to me, I doubt the time between “AI developed that is capable of improving itself” and “AI is applied to improve itself” will differ much between Chinese and Western labs.
So many people believe this “we only need to align a human level researcher” idea now. Leopold basically doesn’t say at all what these automated alignment researchers should be actually doing all day. I feel like a more serious thinker would have at least tried to sketch out what these agents should be working on in alignment or would have started looking into alignment and tried to figure out how hard it is,
If you’re worried at all about alignment then China is clearly worse than the US (not because of anything the Trump admin is doing, but mostly because of the AI industry emerging in large part from the AI safety community). Their labs are producing models with terrible prosaic safety standards and seem to have no interest whatsoever in more significant alignment, and the government seems to share these views. I also think that the Trump admin’s policy, while hamfisted and possibly ill-intended, might be good for reducing risks from cyber and bio-attacks in the short/medium term.
Obviously this mostly makes sense if you think alignment won’t happen by default but is tractable, if you think it’s easy or virtually impossible this doesn’t matter.
I dont really see how it follows any more, there’s no real endgame alignment today either way, and China has expressed concerns as well, they just dont care about today’s models’ alignment as much.
The Trump administration, which seems to be in control more and more than the labs themselves are much more likely to force alignment to themselves, or at best America, and use their advantage to do what they’ve been doing for the last few years—America first at the expense of everyone. China probably ‘just’ forces me to be pro-CCP, and at least allows me to use their models as much as anyone else.
Either seem unlikely to produce true alignment, but if alignment is on the easy side, only one side is truly antagonistic against me and has shown theyll use their advantage to crush me.
Like I said if alignment is impossibly hard then it doesn’t matter whether the US or China win the race (we’re all going to die anyway). If it’s easy then it’s irrelevant and whether you prefer the US or China to win will just come down to which you like more.
If alignment is tractable I suspect we will only solve it by experimenting with less capable models to discover failure modes and by making sure that pre-ASI models are reasonably aligned, since they will probably play a major role in ASI development.
>Like I said if alignment is impossibly hard then it doesn’t matter whether the US or China win the race (we’re all going to die anyway)
Not true, as China gives me more—the ability to use their models—in the mean time.
>If alignment is tractable
It can still be tractable in the sense that we dont die, while making someone the sole benefactor—the US is heavily signaling that’s what they’ll do if they are able to.
You think the main argument in favor of the US is that they have more alignment researchers at their labs? I think in the past people argued that “liberal democracy” must win (Leopold). Today this doesn’t even come up in peoples minds?
The biggest outcome of that is perhaps how immediately they all converged on RSI.
Leopold is pretty clearly on the ‘alignment is relatively easy’ side (“I am not a doomer. Misaligned superintelligence is probably not the biggest AI risk...What I want to do is explain what I see as the “default” plan for how we’ll muddle through, and why I’m optimistic.”). As I noted if you think x risks from misalignment are very low then there is no reason to favour the western approach on alignment (essentially, caring about it at all) since both approaches will work (and if you think it’s extremely high then there is likewise no reason since both approaches will fail).
RSI seems like an obvious idea to me, I doubt the time between “AI developed that is capable of improving itself” and “AI is applied to improve itself” will differ much between Chinese and Western labs.
So many people believe this “we only need to align a human level researcher” idea now. Leopold basically doesn’t say at all what these automated alignment researchers should be actually doing all day. I feel like a more serious thinker would have at least tried to sketch out what these agents should be working on in alignment or would have started looking into alignment and tried to figure out how hard it is,