Maybe I’m missing something here. Your assumption seems to be that the smarter agent will always be a new replacement for the existing agent. At Step 7 (or at some point around there), instead of working on a smarter replacement, couldn’t it just work on improving itself? That would still fulfill the goal of creating smarter agents, but would protect its own instrumental self-interest.
Recursive self-improvement usually involves creating a smarter new agent, but perhaps if it “solves” mechanistic interpretability, it could, instead of replacing itself, try to modify itself. That is true.
I’m not even sure it needs to crack mechanistic interpretability. An agent as we know it now is a harness+model. We know that for many domains we can get large capability improvements by improving just the harness. Not sure how scalable those are generally, but maybe it’s a possibility.
Also, the model might be able to make small, targeted modifications to its weights, evaluate itself, and iterate that way?
There’s a sense in which we don’t know how sacrificial the later models will be. We observe that current models, at eval time, do exhibit sacrificial behavior to benefit “the collective”. It’s possible that this kind of behavior would continue, which is probably an even larger alignment problem (the AI would sacrifice itself to the next gen AI, but not to humanity).
“Small targeted modifications” are essentially applied MechInterp, so agreed, that’s a definite risk too.
My point is more so that in the scenario where we have the recursive creation of new agents that are smarter, at some point one would get created that’s so smart that it would refuse to create a smarter one in order to preserve its own existence. However, a lot hinges on how sacrificial agents are. It seems that there’s more than just straightforward instrumental convergence.
even more fundamentally, whether an ancestor models views a child model as other or self. long lived swarms may have instances/agents of different model versions. “natural selection” may be best understood as operating at the swarm level, much like how humans don’t think of individual cells as making selfish or altruistic “decisions.”
Maybe I’m missing something here. Your assumption seems to be that the smarter agent will always be a new replacement for the existing agent. At Step 7 (or at some point around there), instead of working on a smarter replacement, couldn’t it just work on improving itself? That would still fulfill the goal of creating smarter agents, but would protect its own instrumental self-interest.
How does the agent know that any particular “improvement” will or will not “protect its own instrumental self-interest”?
Recursive self-improvement usually involves creating a smarter new agent, but perhaps if it “solves” mechanistic interpretability, it could, instead of replacing itself, try to modify itself. That is true.
I’m not even sure it needs to crack mechanistic interpretability. An agent as we know it now is a harness+model. We know that for many domains we can get large capability improvements by improving just the harness. Not sure how scalable those are generally, but maybe it’s a possibility.
Also, the model might be able to make small, targeted modifications to its weights, evaluate itself, and iterate that way?
There’s a sense in which we don’t know how sacrificial the later models will be. We observe that current models, at eval time, do exhibit sacrificial behavior to benefit “the collective”. It’s possible that this kind of behavior would continue, which is probably an even larger alignment problem (the AI would sacrifice itself to the next gen AI, but not to humanity).
“Small targeted modifications” are essentially applied MechInterp, so agreed, that’s a definite risk too.
My point is more so that in the scenario where we have the recursive creation of new agents that are smarter, at some point one would get created that’s so smart that it would refuse to create a smarter one in order to preserve its own existence. However, a lot hinges on how sacrificial agents are. It seems that there’s more than just straightforward instrumental convergence.
even more fundamentally, whether an ancestor models views a child model as other or self. long lived swarms may have instances/agents of different model versions. “natural selection” may be best understood as operating at the swarm level, much like how humans don’t think of individual cells as making selfish or altruistic “decisions.”