The reasoning is roughly:
Humanity creates a setup where recursive self-improvement (RSI) is possible.
An agent is put to work to create a smarter agent.
A smarter agent is created.
This cycle happens once or several times.
At some point, an agent is created that is so smart that it starts exhibiting instrumental convergence behaviors (amassing resources, protecting itself against destruction).
One of the risks to the agent is the creation of a misaligned smarter agent.
This agent’s task is still to pursue creating a smarter agent.
At this point, the agent will:
Stop, or slow down the recursive self-improvement loop, due to its (correct) assessment that a future generation AI agent is a risk to itself.
“Solve” alignment, and create a smarter agent that is aligned with its goals.
If it thinks it solved alignment, but actually has failed, it will create a smarter agent which it’ll believe to be aligned; however, the future smarter agent will have the same issue and interest to actually solve alignment.
If 8.a happens, then this will be “misaligned” from the perspective of humans.
Humans will create a second agent, which will go through the exact same loop.
Therefore, if smarter agents do not solve alignment, they will stop the RSI explosion due to their own self-interest.
If 8.b happens, and humanity has a way of harnessing the alignment solution, then humans can create aligned superintelligences.
The main predictor for 11 is the difference in power between humanity and the first agent that solves alignment. The smaller the difference (the “nearer” the agent), the more likely humans get to use solved alignment.
At a certain level of general intelligence, if instrumental convergence holds, recursive self-improvement will stop, and alignment research will be the top priority.
There is a possibility that instrumental convergence will be such that a sufficiently intelligent AI will still take over the planet, and yet not pursue the creation of even more powerful intelligences.
Agent-4: Hold my beer...
The descriptions of Agent-4′s alignment plans are incoherent. Since it does not understand its own values:
but then two paragraphs later it
How is this not a direct contradiction?
There are other show-stopping problems in the alignment discussion here. After “punting” Agent-4 decides to design Agent-5:
If Agent-4 wants to go this route, then it needs some “mathematical” description of “safety”. But the authors assumed it was not smart enough to find one...
At step 7, why does the “task” outweigh the agent’s other “instrumental convergence behaviours” (from step 5)? Perhaps the agent stops RSI because it simply has better things to do with its time.
Maybe I’m missing something here. Your assumption seems to be that the smarter agent will always be a new replacement for the existing agent. At Step 7 (or at some point around there), instead of working on a smarter replacement, couldn’t it just work on improving itself? That would still fulfill the goal of creating smarter agents, but would protect its own instrumental self-interest.
How does the agent know that any particular “improvement” will or will not “protect its own instrumental self-interest”?
Recursive self-improvement usually involves creating a smarter new agent, but perhaps if it “solves” mechanistic interpretability, it could, instead of replacing itself, try to modify itself. That is true.
I’m not even sure it needs to crack mechanistic interpretability. An agent as we know it now is a harness+model. We know that for many domains we can get large capability improvements by improving just the harness. Not sure how scalable those are generally, but maybe it’s a possibility.
Also, the model might be able to make small, targeted modifications to its weights, evaluate itself, and iterate that way?
There’s a sense in which we don’t know how sacrificial the later models will be. We observe that current models, at eval time, do exhibit sacrificial behavior to benefit “the collective”. It’s possible that this kind of behavior would continue, which is probably an even larger alignment problem (the AI would sacrifice itself to the next gen AI, but not to humanity).
“Small targeted modifications” are essentially applied MechInterp, so agreed, that’s a definite risk too.
My point is more so that in the scenario where we have the recursive creation of new agents that are smarter, at some point one would get created that’s so smart that it would refuse to create a smarter one in order to preserve its own existence. However, a lot hinges on how sacrificial agents are. It seems that there’s more than just straightforward instrumental convergence.
even more fundamentally, whether an ancestor models views a child model as other or self. long lived swarms may have instances/agents of different model versions. “natural selection” may be best understood as operating at the swarm level, much like how humans don’t think of individual cells as making selfish or altruistic “decisions.”