cousin_it and I disagreed in comment sections a couple of times so decided to make that discussion into a post.
As we iterated drafts, I realized that we were making separate points more than disagreeing (although we do still disagree on my central point, how risky it would be to have one person control obedient singleton ASI).
cousin_it was focused on the point “for god’s sake let’s not put the elite in charge of transformative AI; they’ll mistreat everyone” and I was making the point “for god’s sake let’s not give everyone their own takeover-capable AI; one person in charge would be less risky than that”.
Those two points are compatible. We should do neither. I think we agree that it’s better to just not build ASI, or to build value-aligned ASI. If either of those is possible.
The reason I think the discussion is worthwhile even when we’re in agreement that we should do neither, is that either moderately distributed or very concentrated power over obedient ASI seems like the most likely outcome on the current path (if we can even pull off instruction-following alignment successfully). So deciding between those two bad options may be all we get to do.
I’m not even sure I disagree with cousin_it on how powerful AI should get before we restrict public access and prevent proliferation. I certainly don’t want to shut that down for open-source models, even next-gen ones capable of truly dangerous hacking and bioengineering. When proliferation becomes too dangerous is going to be a judgment call, and I won’t be the one to make it—those leading the AI race will. (I think this will probably be the sitting POTUS, and that’s the biggest lever we’re likely to get).
I don’t want to see that happen and I don’t want to be one of the casualties; but I want even more for us to make it to aligned ASI and a staggeringly good future. To me the situation looks pretty dire and we’re going to have to take some chances. The trick is calculating our risks as best we can in advance.
The tradeoff I primarily see is having warning shots big enough to make us take this shit seriously but not big enough to kill us all. Of course, the wrong warning shots like a bioengineered plauge created under AI orders could hurt, making us rush to ASI and get egregiously misaligned ASI… which kills us all.
We need to work faster and harder to analyze those risks, because the decisions may be coming up sooner than we’d like.
I want even more for us to make it to aligned ASI and a staggeringly good future. To me the situation looks pretty dire and we’re going to have to take some chances
I’m on board, but open-weight models with superhuman biohacking skills don’t seem like a chance we have to take? Maybe I missing some second order effect that makes fighting for open-weight models in particular worth the inevitable death toll?
Well I’d rather they not have superhuman biohacking skills! We don’t seem on track to be able to preserve some capabilities and prevent others in open-source models. I guess it’s possible to just eliminate huge chunks of the pretraining data, and I sure hope people do that once we’re hitting truly dangerous capabilities (we’ll get our warning shots from whatever shenanigans aren’t totally prevented). Closed-source models are looking pretty easy to monitor and control through wrappers like Fable’s but those won’t allow any resistance to those controlling the companies (which will btw probably be the government).
Here’s a reflection on this debate process:
cousin_it and I disagreed in comment sections a couple of times so decided to make that discussion into a post.
As we iterated drafts, I realized that we were making separate points more than disagreeing (although we do still disagree on my central point, how risky it would be to have one person control obedient singleton ASI).
cousin_it was focused on the point “for god’s sake let’s not put the elite in charge of transformative AI; they’ll mistreat everyone” and I was making the point “for god’s sake let’s not give everyone their own takeover-capable AI; one person in charge would be less risky than that”.
Those two points are compatible. We should do neither. I think we agree that it’s better to just not build ASI, or to build value-aligned ASI. If either of those is possible.
The reason I think the discussion is worthwhile even when we’re in agreement that we should do neither, is that either moderately distributed or very concentrated power over obedient ASI seems like the most likely outcome on the current path (if we can even pull off instruction-following alignment successfully). So deciding between those two bad options may be all we get to do.
I’m not even sure I disagree with cousin_it on how powerful AI should get before we restrict public access and prevent proliferation. I certainly don’t want to shut that down for open-source models, even next-gen ones capable of truly dangerous hacking and bioengineering. When proliferation becomes too dangerous is going to be a judgment call, and I won’t be the one to make it—those leading the AI race will. (I think this will probably be the sitting POTUS, and that’s the biggest lever we’re likely to get).
Really? doesn’t public availability of open-weight superhuman bioengineering basically guarantee megadeaths? what’s the worthwhile trade-off here?
I don’t want to see that happen and I don’t want to be one of the casualties; but I want even more for us to make it to aligned ASI and a staggeringly good future. To me the situation looks pretty dire and we’re going to have to take some chances. The trick is calculating our risks as best we can in advance.
The tradeoff I primarily see is having warning shots big enough to make us take this shit seriously but not big enough to kill us all. Of course, the wrong warning shots like a bioengineered plauge created under AI orders could hurt, making us rush to ASI and get egregiously misaligned ASI… which kills us all.
We need to work faster and harder to analyze those risks, because the decisions may be coming up sooner than we’d like.
I’m on board, but open-weight models with superhuman biohacking skills don’t seem like a chance we have to take? Maybe I missing some second order effect that makes fighting for open-weight models in particular worth the inevitable death toll?
Well I’d rather they not have superhuman biohacking skills! We don’t seem on track to be able to preserve some capabilities and prevent others in open-source models. I guess it’s possible to just eliminate huge chunks of the pretraining data, and I sure hope people do that once we’re hitting truly dangerous capabilities (we’ll get our warning shots from whatever shenanigans aren’t totally prevented). Closed-source models are looking pretty easy to monitor and control through wrappers like Fable’s but those won’t allow any resistance to those controlling the companies (which will btw probably be the government).
Right, so we have to outlaw them. or have Mythos hack and poison them, or something. but no, because … they are crucial to friendly AGI? (are they?)
or is this just the “we can’t install a traffic light until someone actually dies” thing?
Concentration of power is worse.