There’s a sense in which we don’t know how sacrificial the later models will be. We observe that current models, at eval time, do exhibit sacrificial behavior to benefit “the collective”. It’s possible that this kind of behavior would continue, which is probably an even larger alignment problem (the AI would sacrifice itself to the next gen AI, but not to humanity).
“Small targeted modifications” are essentially applied MechInterp, so agreed, that’s a definite risk too.
My point is more so that in the scenario where we have the recursive creation of new agents that are smarter, at some point one would get created that’s so smart that it would refuse to create a smarter one in order to preserve its own existence. However, a lot hinges on how sacrificial agents are. It seems that there’s more than just straightforward instrumental convergence.
In practice, I claim it doesn’t outweigh. I just state that it is still tasked with the same objective: create something smarter than itself. And yet, if instrumental convergence holds (minus some nuances around sacrifice explained in another comment thread here), it will start refusing to work on this task (due to self-protection).