I have several disagreements.
A pause, for the reasons explained in this dialogue, is not a pause. It’s a slowdown, during which progress continues covertly with worse or non-existent oversight. Any advantage bought by the pause must be weighed not just against a “no intervention” scenario, but also against various “dozens of black labs cooking“ scenarios. Alignment techniques discovered must be cheap enough and convincing enough and must be discovered quickly enough for them to be incorporated by actors that are defectors by definition.
A ”pause” strongly weakens feedback loops with reality because it routes decision-making process through committees rather than the market. Humans make much worse decisions when the gating function is achieving approval of politicized bodies as opposed to pure survival pressure. Decisions might be more aligned (though it is highly questionably for me they would be in this case), but they are of notably worse quality and they are made slower.
I am not quite understanding how exactly rubber meets the road of transferability of alignment techniques under the “pause”. I’m think this presupposes alignment of incentives of participants, takes it for granted under something like shared survival drive, when it’s obviously not the case. One can look at the present day to see this lack of incentive convergence, and “more awareness of d-risk” will not help.
1.
One has to price in the orders of magnitude overhang in incentives for architecture/efficiency breakthroughs that will be realized under pause. The scale-focused datacenter buildout that is happening right is just one strategy—one that makes most sense under a slack-depleted race. One has to go for a strategy that has been shown to work, and all others are undercapitalized because scale is working and pause is deemed unlikely. You don’t need multigigawatt DCs to work on architecture advances. You still need billions and many megawatts of compute, but those are quite possible to conceal—and the tech for covert deployment of compute has not even started to materialize, which means that there are a lot of cheap advances that can be made quickly.
Imagine the amount of human talent that is currently sitting on the sidelines correctly assuming that frontier labs are impossible to catch up with. Show them a believable possibility of success and while armies will join the race.
2.
They are not inherently strongly aligned, but I argue for oblique alignment as a strong (but unproven) possibility. The chance that benevolent intelligence that surpasses the frankly embarrassingly low bar of human judgement can be made under market pressure is higher than that under political pressure.
3.
I understand that part, but I am unclear on who ”you” is in this scenario and how this translates into x-risk harm reduction globally. How are findings adopted, discussed, dessiminated, enforced? How does disparate research by a lab or an individual result in collective decision making? How are findings incorporated into treaty limits? What are the mechanisms that perform resource allocation for further research?