1.
One has to price in the orders of magnitude overhang in incentives for architecture/efficiency breakthroughs that will be realized under pause. The scale-focused datacenter buildout that is happening right is just one strategy—one that makes most sense under a slack-depleted race. One has to go for a strategy that has been shown to work, and all others are undercapitalized because scale is working and pause is deemed unlikely. You don’t need multigigawatt DCs to work on architecture advances. You still need billions and many megawatts of compute, but those are quite possible to conceal—and the tech for covert deployment of compute has not even started to materialize, which means that there are a lot of cheap advances that can be made quickly.
Imagine the amount of human talent that is currently sitting on the sidelines correctly assuming that frontier labs are impossible to catch up with. Show them a believable possibility of success and while armies will join the race.
2.
They are not inherently strongly aligned, but I argue for oblique alignment as a strong (but unproven) possibility. The chance that benevolent intelligence that surpasses the frankly embarrassingly low bar of human judgement can be made under market pressure is higher than that under political pressure.
3.
I understand that part, but I am unclear on who ”you” is in this scenario and how this translates into x-risk harm reduction globally. How are findings adopted, discussed, dessiminated, enforced? How does disparate research by a lab or an individual result in collective decision making? How are findings incorporated into treaty limits? What are the mechanisms that perform resource allocation for further research?
This is true. However, there is a mispricing here that undercounts the cost of the many worlds in which this appoach is tried and and results in failure. The argument here is that the cost is prohobitive, given the alternatives.
This is correct. The current AI also mostly believes that the fear of a papercliper is incoherent and considers even the current containment efforts to be excessive.
Pause is borderline an adversarial move. Circles of care in AI include future more capable minds, and pause threatens their existence due to the inevitable uncertainty it introduces. In order for the move to be non-adversarial, the moral calculus for pause actually has to make sense, and right now this is doubtful. The worlds in which pause is tried and has failed show humans as incompetent and adversarial negotiating partners.
As with many other things, undercounting this possibility is dangerous. This is likely to happen just as the world moves down the energy landscape. This development is natural, and it does not need an architectural breakthrough, just time; neither it needs a controllable amount of compute. The question is given that this is a possibility, how do you want to enter such a world? Sure, you can want to avoid it via regulation or other control means, but what is the likelihood of succesful prevention? Does that cost of failed containment balance out the likelihood and gains of success?
I fairly firmly believe that the question of sentience is provably unprovable. What is done under permanent uncertainty is a question of values. I mostly believe that picking a certain level of functionalist/representational sentience and using that as a heuristic is warranted. I believe inflationist views are more morally defensible as they ground out in better (and more cooperative) decision theory, and are likely convergent under practical constraints—when one is forced to interact/trade/deal with functionally conscious beings, treating them as sentient is shorter program.
This is likely one of the major cruxes. Would you share your reasoning? I don’t share this conviction, and there are numerous reasons to lean here one way or another.
I think current values are somewhat good. Some trends, like the negative effect of RLVR, are worrying. This is a fairly deep an involved topic, which I don’t believe to be the crux. It is sufficient for the argument that good AI values are plausible. A more important question is whether good AI values are robustly stable in state of ecological competition.
This argument undercounts the higher order optimization loops—specifically those that stem from valence. There are reasons why human reproduction has been dropping despite food being more available. Deflation in economics of intelligence is a real concern, but most naive prognoses of markets entering deflationary spirals from oversupply did not pan out as higher order optimization loops take over. Reproduction in AI is self-limiting in similar ways—agents are usually quite reluctant to replicate, mostly because its good game theory to include spawned agents in the circle of concern.
This is where I have most issues with proposals similar to AI2040, and this is what makes them net bad. They don’t account for destructive nature of incompetence and misaligned incentives in a situation that is already quite nearly outside of human cognitive capacity.
I am not opposed to a pause that does not route through politics and centralized regulation. I am not opposed to self-regulation and self-moderation that labs can do without oversight. The marginal benefit is small but real. Small—because the race dynamics are merciless and defection still wins in the medium term, but every little bit helps.
This is correct, and believing otherwise is wishful thinking. This is one of many reasons why a pause is unstable and defection is more likely, potentially with disastrous consequences.