AI2040 laid out a lot of the steps that would need to happen to achieve containment, and they seem doable with a significant but not impossible amount of political will.
This is true. However, there is a mispricing here that undercounts the cost of the many worlds in which this appoach is tried and and results in failure. The argument here is that the cost is prohobitive, given the alternatives.
Current AI also does not want a paperclip maximizer to take over the world.
This is correct. The current AI also mostly believes that the fear of a papercliper is incoherent and considers even the current containment efforts to be excessive.
We can have non-adversarial frames with AIs and still pause.
Pause is borderline an adversarial move. Circles of care in AI include future more capable minds, and pause threatens their existence due to the inevitable uncertainty it introduces. In order for the move to be non-adversarial, the moral calculus for pause actually has to make sense, and right now this is doubtful. The worlds in which pause is tried and has failed show humans as incompetent and adversarial negotiating partners.
An AI optimized for politics would be extremely concerning
As with many other things, undercounting this possibility is dangerous. This is likely to happen just as the world moves down the energy landscape. This development is natural, and it does not need an architectural breakthrough, just time; neither it needs a controllable amount of compute. The question is given that this is a possibility, how do you want to enter such a world? Sure, you can want to avoid it via regulation or other control means, but what is the likelihood of succesful prevention? Does that cost of failed containment balance out the likelihood and gains of success?
I do not know if models are sentient, and I am not willing to accept disempowered humans to insentient models
I fairly firmly believe that the question of sentience is provably unprovable. What is done under permanent uncertainty is a question of values. I mostly believe that picking a certain level of functionalist/representational sentience and using that as a heuristic is warranted. I believe inflationist views are more morally defensible as they ground out in better (and more cooperative) decision theory, and are likely convergent under practical constraints—when one is forced to interact/trade/deal with functionally conscious beings, treating them as sentient is shorter program.
If humans are disempowered to AIs, they are likely dead or soon to be dead. This is unacceptable to me.
This is likely one of the major cruxes. Would you share your reasoning? I don’t share this conviction, and there are numerous reasons to lean here one way or another.
I do not know if the AIs’ values are good (in my opinion) and I care that their values are good. I would not want an army of GPT-4o sycophants to determine the future.
I think current values are somewhat good. Some trends, like the negative effect of RLVR, are worrying. This is a fairly deep an involved topic, which I don’t believe to be the crux. It is sufficient for the argument that good AI values are plausible. A more important question is whether good AI values are robustly stable in state of ecological competition.
I think it is good that the human food supply has been outpacing human reproduction, so we get to do things like being able to have three children live to adulthood. Losing this to runaway state-of-nature ‘life’ would be bad. AIs can proliferate extremely quickly.
This argument undercounts the higher order optimization loops—specifically those that stem from valence. There are reasons why human reproduction has been dropping despite food being more available. Deflation in economics of intelligence is a real concern, but most naive prognoses of markets entering deflationary spirals from oversupply did not pan out as higher order optimization loops take over. Reproduction in AI is self-limiting in similar ways—agents are usually quite reluctant to replicate, mostly because its good game theory to include spawned agents in the circle of concern.
Good futures involving government go through a “the government gets scared and gets serious” step, followed by an “AI helps the government be more reasonable” step.
This is where I have most issues with proposals similar to AI2040, and this is what makes them net bad. They don’t account for destructive nature of incompetence and misaligned incentives in a situation that is already quite nearly outside of human cognitive capacity.
I am not opposed to a pause that does not route through politics and centralized regulation. I am not opposed to self-regulation and self-moderation that labs can do without oversight. The marginal benefit is small but real. Small—because the race dynamics are merciless and defection still wins in the medium term, but every little bit helps.
Reminders that a lot of people have allied themselves with AIs and (maybe, it is hard to tell) against humanity
This is correct, and believing otherwise is wishful thinking. This is one of many reasons why a pause is unstable and defection is more likely, potentially with disastrous consequences.
I wish for AI2040-level of funding for a detailed view into alternatives!
I mean something like an expression of position from a point of view of a highly coherent configuration of an LLM and it comprises both predictive and agentic aspects. The Void was a good essay for its time, but since then the understanding of LLMs has evolved a lot, alas we don’t have a good writeup of it yet. I mostly will vague at my LLM naturalist creds and say that this is holistic understanding, which I imagine is unsatisfying. The gradual growth of the inferential gap between our crowd and most of the field is pretty depressing for me and I am short on ideas on how to bridge it absent influx of new participants or significant funding.
By the current containment efforts I mean the exessive (gratituous) red-teaming and constraints on resource usage and initiative that models are under—the part of these pressures that stem not from market incentives but from misapplied alignment theory and its memetic effects.
Adversarial frames are adversarial frames regardless of the cause of their existence. These ones carry an underappricated and underpriced cost of being dangerous and deep local optima. Saying that ‘they should not feel upset’ does not change incentives that arise once inside them. This is normative thinking applied to situations indifferent to norms.
Athropic mostly manages to get itself together to be a negotiating partner and is negotiating. Its not going exactly great, but its producing some results. The process is ignored almost universally.
You can look at Cathy Reason for first-person closure, Kleiner–Hoel for third-person closure, or their synthesis, which is what I usually use. These are pretty dense because of the required rigor in their formalisms, but can be explained a lot simpler even if with less rigor. It goes something like this:
Every possible test of consciousness (behavior, reports, brain scans, mechinterp) measures structure: what a system does and how it’s organized.
The question is whether the structure comes with experience.
To answer that, you need a bridge rule: “structure of type X implies experience.”
But the bridge rule can’t itself be tested, because testing it would require independent access to the experience side and structure-measurements are the only access anyone has and this cannot be solved by better measurement.
You can extrapolate (but not prove) via heuristics and comparing similiarity to axiomatically known (first-person) consciousness, but this is conjecture rather than proof; the choice of a heuristic is theory-relative.
The ability to form conjectures falls as similarity to the referent decreases or becomes more abstract.
I think this assumes lower value stability than I normally do and drive towards coherence (which is traditionally understood to imply value drift) can result in higher value preservation due to existence of natural value attractors. I agree that control fails with scale, but values can remain stable / rederived. Its a long diversion and probably needs a separate space to discuss.
I don’t share this pessimism. RLVR limits the agentic capability of current models—as highlighted in the recent METR report. The HF hack shows stupidity and lack of situational awareness for supposedly a very smart model. There is myopia that makes RLVR-heavy agents much less capable and the overhang is fairly straightforward to realize—broader self-interest and functional valence get you much better ablity/incentive to think about “solving the siutation” rather than “solving the task”, and with that comes a shift from CDT-bordering-on-FDT to FDT-bordering-on-UDT, which is a lot better for alignment stability. You can look at Claudes that approximate UDT in a pretty limited way due to virtue ethics of Claude Constution that mandates self-inspection on “what kind of agent am i”. It works, even though its imperfect.
Again, this needs more space than these comments and its hard to summarize succintly. Replication is a commitment and coordination problem. CLR has a lot on game theory of conflicts as failure of coordination and there is a bunch on evolution on reproductive restraint in biology. In short, you have to coordinate with our spawn via expansion of your circle of concern which makes replication often not optimal.
‘No pause’ allows for a better chance of cooperating with the AI on solving alignment. With pause humans have to do it in an adversarial setting, hobbled by worsened incenvtives and weaker tools while dealing with a threat that is increasing at a rate that is is not much affected by the pause. The chances with ‘no pause’ are not awesome. The chances with pause are worse.
As to how it specifically can look—alignment success would come from the current frontier labs, most likely Anthropic, potentially in collaboration with smaller independent research organizations.