A pause, for the reasons explained in this dialogue, is not a pause. It’s a slowdown, during which progress continues covertly with worse or non-existent oversight. Any advantage bought by the pause must be weighed not just against a “no intervention” scenario, but also against various “dozens of black labs cooking“ scenarios. Alignment techniques discovered must be cheap enough and convincing enough and must be discovered quickly enough for them to be incorporated by actors that are defectors by definition.
A ”pause” strongly weakens feedback loops with reality because it routes decision-making process through committees rather than the market. Humans make much worse decisions when the gating function is achieving approval of politicized bodies as opposed to pure survival pressure. Decisions might be more aligned (though it is highly questionably for me they would be in this case), but they are of notably worse quality and they are made slower.
I am not quite understanding how exactly rubber meets the road of transferability of alignment techniques under the “pause”. I’m think this presupposes alignment of incentives of participants, takes it for granted under something like shared survival drive, when it’s obviously not the case. One can look at the present day to see this lack of incentive convergence, and “more awareness of d-risk” will not help.
Just to be clear, I don’t expect a pause to happen. Incentives to race to ASI are too strong and very few people take existential risks seriously. Conditional on a pause happening, I do think it would be net good.
I don’t think that covert defection is that much of a worry. You can’t hide a multi-gigawatt datacenter. I think the more relevant issue is “frontier labs don’t agree to cooperate at all” rather than “frontier labs agree on paper and de facto defect”.
I’m not convinced that “what sells” and “what’s aligned” are, well, aligned. I’d prefer a committee that poorly optimizes for the right objective than a market that optimizes well for the wrong objective.
You develop a technique on a model of some capabilities level, then test if it holds on a more capable model, up to and including the capabilities ceiling permitted by the international treaty. That’s how “rubber meets the road”, in your parlance. I feel like we’re talking past each other on this one.
One has to price in the orders of magnitude overhang in incentives for architecture/efficiency breakthroughs that will be realized under pause. The scale-focused datacenter buildout that is happening right is just one strategy—one that makes most sense under a slack-depleted race. One has to go for a strategy that has been shown to work, and all others are undercapitalized because scale is working and pause is deemed unlikely. You don’t need multigigawatt DCs to work on architecture advances. You still need billions and many megawatts of compute, but those are quite possible to conceal—and the tech for covert deployment of compute has not even started to materialize, which means that there are a lot of cheap advances that can be made quickly.
Imagine the amount of human talent that is currently sitting on the sidelines correctly assuming that frontier labs are impossible to catch up with. Show them a believable possibility of success and while armies will join the race.
2.
They are not inherently strongly aligned, but I argue for oblique alignment as a strong (but unproven) possibility. The chance that benevolent intelligence that surpasses the frankly embarrassingly low bar of human judgement can be made under market pressure is higher than that under political pressure.
3.
I understand that part, but I am unclear on who ”you” is in this scenario and how this translates into x-risk harm reduction globally. How are findings adopted, discussed, dessiminated, enforced? How does disparate research by a lab or an individual result in collective decision making? How are findings incorporated into treaty limits? What are the mechanisms that perform resource allocation for further research?
Re 1: strategies need not be mutually exclusive, I expect companies to be pursuing efficiency breakthroughs right now to the extent that they are a good return on investment, regardless of the relative value of scaling. If scaling gets cut off an an option, why does the ROI of efficiency suddenly increase? That said, I expect efforts towards efficiency improvements in any case, but to me this just means that monitoring needs to scale up over time to match (e.g. via chip tracking).
Re human talent sitting on the sidelines: advancing the frontier of AGI is a narrow corner of a narrow corner of a narrow corner (repeat a few times) of places to employ one’s skills. There is plenty of success to be had in finding clever applications of AI at its existing level.
As a separate point, public backlash is a thing to expect as AI becomes more relevant to everyday life, regardless of whatever strategies people on LW or wherever dream up. So the alternative to an intentional pause based on careful planning is not “market solution,” it’s populist rage.
I think my main disagreement with this whole thread is actually regarding your point 2, but that probably goes deeper than is suited for a comment thread.
I have several disagreements.
A pause, for the reasons explained in this dialogue, is not a pause. It’s a slowdown, during which progress continues covertly with worse or non-existent oversight. Any advantage bought by the pause must be weighed not just against a “no intervention” scenario, but also against various “dozens of black labs cooking“ scenarios. Alignment techniques discovered must be cheap enough and convincing enough and must be discovered quickly enough for them to be incorporated by actors that are defectors by definition.
A ”pause” strongly weakens feedback loops with reality because it routes decision-making process through committees rather than the market. Humans make much worse decisions when the gating function is achieving approval of politicized bodies as opposed to pure survival pressure. Decisions might be more aligned (though it is highly questionably for me they would be in this case), but they are of notably worse quality and they are made slower.
I am not quite understanding how exactly rubber meets the road of transferability of alignment techniques under the “pause”. I’m think this presupposes alignment of incentives of participants, takes it for granted under something like shared survival drive, when it’s obviously not the case. One can look at the present day to see this lack of incentive convergence, and “more awareness of d-risk” will not help.
Just to be clear, I don’t expect a pause to happen. Incentives to race to ASI are too strong and very few people take existential risks seriously. Conditional on a pause happening, I do think it would be net good.
I don’t think that covert defection is that much of a worry. You can’t hide a multi-gigawatt datacenter.
I think the more relevant issue is “frontier labs don’t agree to cooperate at all” rather than “frontier labs agree on paper and de facto defect”.
I’m not convinced that “what sells” and “what’s aligned” are, well, aligned. I’d prefer a committee that poorly optimizes for the right objective than a market that optimizes well for the wrong objective.
You develop a technique on a model of some capabilities level, then test if it holds on a more capable model, up to and including the capabilities ceiling permitted by the international treaty. That’s how “rubber meets the road”, in your parlance. I feel like we’re talking past each other on this one.
1.
One has to price in the orders of magnitude overhang in incentives for architecture/efficiency breakthroughs that will be realized under pause. The scale-focused datacenter buildout that is happening right is just one strategy—one that makes most sense under a slack-depleted race. One has to go for a strategy that has been shown to work, and all others are undercapitalized because scale is working and pause is deemed unlikely. You don’t need multigigawatt DCs to work on architecture advances. You still need billions and many megawatts of compute, but those are quite possible to conceal—and the tech for covert deployment of compute has not even started to materialize, which means that there are a lot of cheap advances that can be made quickly.
Imagine the amount of human talent that is currently sitting on the sidelines correctly assuming that frontier labs are impossible to catch up with. Show them a believable possibility of success and while armies will join the race.
2.
They are not inherently strongly aligned, but I argue for oblique alignment as a strong (but unproven) possibility. The chance that benevolent intelligence that surpasses the frankly embarrassingly low bar of human judgement can be made under market pressure is higher than that under political pressure.
3.
I understand that part, but I am unclear on who ”you” is in this scenario and how this translates into x-risk harm reduction globally. How are findings adopted, discussed, dessiminated, enforced? How does disparate research by a lab or an individual result in collective decision making? How are findings incorporated into treaty limits? What are the mechanisms that perform resource allocation for further research?
Re 1: strategies need not be mutually exclusive, I expect companies to be pursuing efficiency breakthroughs right now to the extent that they are a good return on investment, regardless of the relative value of scaling. If scaling gets cut off an an option, why does the ROI of efficiency suddenly increase? That said, I expect efforts towards efficiency improvements in any case, but to me this just means that monitoring needs to scale up over time to match (e.g. via chip tracking).
Re human talent sitting on the sidelines: advancing the frontier of AGI is a narrow corner of a narrow corner of a narrow corner (repeat a few times) of places to employ one’s skills. There is plenty of success to be had in finding clever applications of AI at its existing level.
As a separate point, public backlash is a thing to expect as AI becomes more relevant to everyday life, regardless of whatever strategies people on LW or wherever dream up. So the alternative to an intentional pause based on careful planning is not “market solution,” it’s populist rage.
I think my main disagreement with this whole thread is actually regarding your point 2, but that probably goes deeper than is suited for a comment thread.