Thanks for reading our thing & for this thoughtful critique. You’ve said a lot of things, and I’m writing quickly late at night so I’ll be brief I’m afraid.
Re 1: We debated amongst ourselves how much structure to impose on the regulations. I think the way we did it is best. (a) If you have a better idea for a specific international governance structure that would be better than realpolitik, you are welcome to propose it, and we’re glad to accept it as an improvement on our current Plan A if it works. We might put more effort into thinking of one later. (b) Even the relatively structureless thing we proposed—research transparency + MACD but otherwise the nations of the world just have to muddle through and handle things on a case by case basis—is a significant improvement over the status quo, for reasons we’ve articulated in the piece. You seem to disagree with this but I don’t see why.
Re 2: Yeah I mean if China ends up catching up, or not realizing that a gap of a few months could be fatal, then that significantly undermines our argument for why they’d want a deal. OTOH it would make it more plausible that the US would want a deal. In general the party falling behind will have more reason to want a deal than the party in the lead. It seems like we have a disagreement about whether China will catch up, but it’s probably not a massive one (neither of us are confident) right?
Re 3: Do you think the surveillance possibilities are less bad in Plan D, C, or B? I would argue that the surveillance possibilities are even worse in those worlds than in Plan A. For reasons we’ve discussed (greater number of frontier AI companies, greater number of countries with frontier AI, vastly more visibility of the public into how AIs are trained and regulated, making it harder for corporations and governments to abuse their power over AI) Only Plan S can seriously claim to be less scary viz-a-viz surveillance. If there’s some plan different from all of the above that you prefer, can you sketch what it is?
Fwiw Plan A would still work OK without the open-weights ban and with a much larger fraction of inference ZDR; would those modifications make you happy? (i.e. it would still be better than the so-far-proposed alternatives IMO) I totally take the point that the US and China could start to do Plan A but then switch to a worse, more authoritarian version that e.g. doesn’t have ZDR, but isn’t that also an argument against the other plans on offer too? Do you have a proposal for how to stop governments from being authoritarian? Plan A does something at least—it creates conditions so that more companies, across more countries, have frontier AI, and so that the public has more visibility into it all.
We really tried to make Plan A reduce concentration of power risks. One of the things that saddens me about the reaction by various people online (including some you approvingly link to) is that they seem to have come in with this prejudice that we AI safety people only care about AI takeover and are planning to concentrate power hugely in the service of our perceived greater good of stopping the misaligned AIs. This stereotype has some basis in reality—various past proposals by various AI safety people were basically like that, and I’d argue that the current plans of Anthropic and OpenAI are basically that or worse—but we really truly do care about concentration of power and we thought a lot about how to prevent it, and I hope it came across in the final product.
We might put more effort into thinking of one later. (b) Even the relatively structureless thing we proposed—research transparency + MACD but otherwise the nations of the world just have to muddle through and handle things on a case by case basis—is a significant improvement over the status quo, for reasons we’ve articulated in the piece. You seem to disagree with this but I don’t see why.
(1) Part of my objection is that I anticipate the structurelessness of it to be an obstacle to its adoption.
Like this is what a senior decision-maker in Bejing or Washington will be contemplating: They’re going to put their economy at the mercy of their greatest geopolitical rival. That is, after this deal, it will become relatively trivial for either the US or China to cripple the economy of the other, albeit at the cost of their own economy being subsequently crippled. Of course, you could say that before making the deal they were at risk of being taken over by some AI, but this risk was diffuse and uncertain; the risk that they’re signing up for now is concrete and definite.
And this lever of destruction could be used, of course, for reasons other than to stop an AI takeoff, and decision-makers in both countries will be acutely aware of this. Suppose China decides that, if it cripples everyone’s compute, it would gain a vast differential advantage because its economy depends less on GPUS than the USA’s economy depends: then it would be in China’s advantage to get in a MACD situation, then to trigger it, and subsequently dominate the US; and a US political figure, anticipating this, would object to MACD. Or suppose that China thinks, “Hrm, the US is an unreliable actor, and a quick and bloodless MACD might be triggered by, for instance, a senile or unstable US President, of which the US recently has had a fair number.” Thus, because China doesn’t want to trigger MACD for reasons of random shit in the future, it would object to it. And so on and so forth.
What makes this worse is that part of what makes nuclear MAD a plausible means of peace is that there’s a clear signal of “Have the nukes been launched,” while there isn’t as clear signal of danger in the case of AI takeoff; and a rational decision-maker, seeing this, will update downwards about whether MACD will be triggered for reasons relating to AI takeoff and upwards about whether MACD will be triggered for some other random shit.
That is, part of what makes nuclear MAD work is that there are radar stations in Siberia and Greenland and Canada, which can detect ICBMs and bombers that have been launched. The US knows that it could detect things being launched, and Russia knows that the US knows, and the US knows that Russia knows that the US knows, and so forth. Imagine if, by contrast, the sign for “the nukes have been launched,” was that panel of experts, notorious for disagreeing among themselves, came to a consensus that the nukes had been launched or might have been launched. If this were so, then nuclear MAD would be much less effective as a game-theoretic means of peace. But of course this is the situation that we’re in with regards to AI.
(2) Part of my confusion I just don’t know what parts of the scenario are predictions and which ones are hopes. Like in the response to Thane:
A consortium of multiple governments—some of which are actual democracies thanks to the transparency requirements which help prevent AI-assisted executive power grabs—is way less bad than a single global dictator, for example.
In general, I’m dubious whether democracies other than the US (is the US to be an actual democracy?) to have any decision-making clout in the Consortium, because China would object to the possibility of being outvoted. But like, I don’t know how much of “Consortium influenced by many democracies” is part of the prediction of what (“transparency + MACD”) gets you, given those two goals; or if “Consortium influenced by many democracies” is maybe a bonus that we might get after setting up the structure of “transparency + MACD,” but a bonus that we’re unlikely to get.
I totally take the point that the US and China could start to do Plan A but then switch to a worse, more authoritarian version that e.g. doesn’t have ZDR, but isn’t that also an argument against the other plans on offer too? Do you have a proposal for how to stop governments from being authoritarian?
FWIW, I don’t think anything short of Plan S can avoid concentration of power. Roughly speaking, for any given Plan X that navigates the AI risk while avoiding concentration of power, there is a neighboring Plan X* that is the same except that it concentrates power in the hands of the entities implementing the plan. If a consortium of companies and governments is prompted to implement Plan X, it would always be rational for them to do Plan X* instead. “Plan A” and “Plan A but without ZDR” is just one specific example pair.
Like, AI 2040 argues:
In short, the answer is that anyone concerned about loss of control should think Plan A is an improvement, along with anyone concerned about concentration of power—except for the people in whom the power would concentrate by default.
For these reasons, we expect strong opposition to Plan A from the leading AI companies. We expect them to rationalize arguments for why the deal is bad and why instead what’s best for America and humanity is a different strategy that just so happens to allow them to continue accumulating massive amounts of power.
China, by contrast, is an example of an actor in whom power would not concentrate by default.
Sure, suppose that’s true. Why would China agree to Plan A specifically, instead of Plan A but no ZDR?
This is IMO roughly the same failure mode as “if we have multipolar takeoff with several mutually misaligned ASIs, some of them would ally with humans and so humans would survive”. But no, the entities actually holding the world-changing power in the now can negotiate among themselves to screw everyone else out of having any power in the future.
It is theoretically possible to insist on anti-authoritarian implementations if the public is very aware of what’s happening and screens politicians and policies for that very strongly. But I am really, really pessimistic about that, given the current “vote for the least worst guy” US paradigm. However politically infeasible Plan S may seem, it seems more feasible than this.
I dunno, maybe there are some weaknesses in this argument and some way to design a plan such that there aren’t neighboring “except we also take all the power” plan variants; plans where the plan-implementers structurally can’t collide. I don’t currently see it, though.
There are degrees of concentration of power. A consortium of multiple governments—some of which are actual democracies thanks to the transparency requirements which help prevent AI-assisted executive power grabs—is way less bad than a single global dictator, for example.
I agree though that AGI and RSI are technologies that inherently concentrate power by default. It’s going to be really hard to resist that innate tendency. But I think Plan A does basically the best we can of the available plans so far, except maybe Plan S.
Without the open-weights ban I would expect terrorists engineering pandemics. As for the ZDR policy I would guess that data is not to be retained if it comes from a weak AI or if a classifier decides that the data is not related to hazardous topics in a manner similar to deciding not to reroute queries from Fable 5 to Opus 4.8.
On the other hand, I wonder if concentration of power can’t actually be prevented even by perfect epistemics and coordination of the humans.
Yes, it would probably lead to terrorists engineering pandemics eventually. However, maybe that’s OK. Maybe the d/acc hardening we describe in the early 2030′s of our scenario would be enough to prevent the worst outcomes here, long enough for the robot economy boom to make things even safer by rendering the biosphere unnecessary.
Thanks for reading our thing & for this thoughtful critique. You’ve said a lot of things, and I’m writing quickly late at night so I’ll be brief I’m afraid.
Re 1: We debated amongst ourselves how much structure to impose on the regulations. I think the way we did it is best. (a) If you have a better idea for a specific international governance structure that would be better than realpolitik, you are welcome to propose it, and we’re glad to accept it as an improvement on our current Plan A if it works. We might put more effort into thinking of one later. (b) Even the relatively structureless thing we proposed—research transparency + MACD but otherwise the nations of the world just have to muddle through and handle things on a case by case basis—is a significant improvement over the status quo, for reasons we’ve articulated in the piece. You seem to disagree with this but I don’t see why.
Re 2: Yeah I mean if China ends up catching up, or not realizing that a gap of a few months could be fatal, then that significantly undermines our argument for why they’d want a deal. OTOH it would make it more plausible that the US would want a deal. In general the party falling behind will have more reason to want a deal than the party in the lead. It seems like we have a disagreement about whether China will catch up, but it’s probably not a massive one (neither of us are confident) right?
Re 3: Do you think the surveillance possibilities are less bad in Plan D, C, or B? I would argue that the surveillance possibilities are even worse in those worlds than in Plan A. For reasons we’ve discussed (greater number of frontier AI companies, greater number of countries with frontier AI, vastly more visibility of the public into how AIs are trained and regulated, making it harder for corporations and governments to abuse their power over AI) Only Plan S can seriously claim to be less scary viz-a-viz surveillance. If there’s some plan different from all of the above that you prefer, can you sketch what it is?
Fwiw Plan A would still work OK without the open-weights ban and with a much larger fraction of inference ZDR; would those modifications make you happy? (i.e. it would still be better than the so-far-proposed alternatives IMO) I totally take the point that the US and China could start to do Plan A but then switch to a worse, more authoritarian version that e.g. doesn’t have ZDR, but isn’t that also an argument against the other plans on offer too? Do you have a proposal for how to stop governments from being authoritarian? Plan A does something at least—it creates conditions so that more companies, across more countries, have frontier AI, and so that the public has more visibility into it all.
We really tried to make Plan A reduce concentration of power risks. One of the things that saddens me about the reaction by various people online (including some you approvingly link to) is that they seem to have come in with this prejudice that we AI safety people only care about AI takeover and are planning to concentrate power hugely in the service of our perceived greater good of stopping the misaligned AIs. This stereotype has some basis in reality—various past proposals by various AI safety people were basically like that, and I’d argue that the current plans of Anthropic and OpenAI are basically that or worse—but we really truly do care about concentration of power and we thought a lot about how to prevent it, and I hope it came across in the final product.
(1) Part of my objection is that I anticipate the structurelessness of it to be an obstacle to its adoption.
Like this is what a senior decision-maker in Bejing or Washington will be contemplating: They’re going to put their economy at the mercy of their greatest geopolitical rival. That is, after this deal, it will become relatively trivial for either the US or China to cripple the economy of the other, albeit at the cost of their own economy being subsequently crippled. Of course, you could say that before making the deal they were at risk of being taken over by some AI, but this risk was diffuse and uncertain; the risk that they’re signing up for now is concrete and definite.
And this lever of destruction could be used, of course, for reasons other than to stop an AI takeoff, and decision-makers in both countries will be acutely aware of this. Suppose China decides that, if it cripples everyone’s compute, it would gain a vast differential advantage because its economy depends less on GPUS than the USA’s economy depends: then it would be in China’s advantage to get in a MACD situation, then to trigger it, and subsequently dominate the US; and a US political figure, anticipating this, would object to MACD. Or suppose that China thinks, “Hrm, the US is an unreliable actor, and a quick and bloodless MACD might be triggered by, for instance, a senile or unstable US President, of which the US recently has had a fair number.” Thus, because China doesn’t want to trigger MACD for reasons of random shit in the future, it would object to it. And so on and so forth.
What makes this worse is that part of what makes nuclear MAD a plausible means of peace is that there’s a clear signal of “Have the nukes been launched,” while there isn’t as clear signal of danger in the case of AI takeoff; and a rational decision-maker, seeing this, will update downwards about whether MACD will be triggered for reasons relating to AI takeoff and upwards about whether MACD will be triggered for some other random shit.
That is, part of what makes nuclear MAD work is that there are radar stations in Siberia and Greenland and Canada, which can detect ICBMs and bombers that have been launched. The US knows that it could detect things being launched, and Russia knows that the US knows, and the US knows that Russia knows that the US knows, and so forth. Imagine if, by contrast, the sign for “the nukes have been launched,” was that panel of experts, notorious for disagreeing among themselves, came to a consensus that the nukes had been launched or might have been launched. If this were so, then nuclear MAD would be much less effective as a game-theoretic means of peace. But of course this is the situation that we’re in with regards to AI.
(2) Part of my confusion I just don’t know what parts of the scenario are predictions and which ones are hopes. Like in the response to Thane:
In general, I’m dubious whether democracies other than the US (is the US to be an actual democracy?) to have any decision-making clout in the Consortium, because China would object to the possibility of being outvoted. But like, I don’t know how much of “Consortium influenced by many democracies” is part of the prediction of what (“transparency + MACD”) gets you, given those two goals; or if “Consortium influenced by many democracies” is maybe a bonus that we might get after setting up the structure of “transparency + MACD,” but a bonus that we’re unlikely to get.
__ Re. 2: Yeah checks out
FWIW, I don’t think anything short of Plan S can avoid concentration of power. Roughly speaking, for any given Plan X that navigates the AI risk while avoiding concentration of power, there is a neighboring Plan X* that is the same except that it concentrates power in the hands of the entities implementing the plan. If a consortium of companies and governments is prompted to implement Plan X, it would always be rational for them to do Plan X* instead. “Plan A” and “Plan A but without ZDR” is just one specific example pair.
Like, AI 2040 argues:
Sure, suppose that’s true. Why would China agree to Plan A specifically, instead of Plan A but no ZDR?
This is IMO roughly the same failure mode as “if we have multipolar takeoff with several mutually misaligned ASIs, some of them would ally with humans and so humans would survive”. But no, the entities actually holding the world-changing power in the now can negotiate among themselves to screw everyone else out of having any power in the future.
It is theoretically possible to insist on anti-authoritarian implementations if the public is very aware of what’s happening and screens politicians and policies for that very strongly. But I am really, really pessimistic about that, given the current “vote for the least worst guy” US paradigm. However politically infeasible Plan S may seem, it seems more feasible than this.
I dunno, maybe there are some weaknesses in this argument and some way to design a plan such that there aren’t neighboring “except we also take all the power” plan variants; plans where the plan-implementers structurally can’t collide. I don’t currently see it, though.
There are degrees of concentration of power. A consortium of multiple governments—some of which are actual democracies thanks to the transparency requirements which help prevent AI-assisted executive power grabs—is way less bad than a single global dictator, for example.
I agree though that AGI and RSI are technologies that inherently concentrate power by default. It’s going to be really hard to resist that innate tendency. But I think Plan A does basically the best we can of the available plans so far, except maybe Plan S.
Without the open-weights ban I would expect terrorists engineering pandemics. As for the ZDR policy I would guess that data is not to be retained if it comes from a weak AI or if a classifier decides that the data is not related to hazardous topics in a manner similar to deciding not to reroute queries from Fable 5 to Opus 4.8.
On the other hand, I wonder if concentration of power can’t actually be prevented even by perfect epistemics and coordination of the humans.
Yes, it would probably lead to terrorists engineering pandemics eventually. However, maybe that’s OK. Maybe the d/acc hardening we describe in the early 2030′s of our scenario would be enough to prevent the worst outcomes here, long enough for the robot economy boom to make things even safer by rendering the biosphere unnecessary.