Partially agree.
If AI development is more insight-driven than compute-driven, then there is more room for sudden progress that gains a decisive advantage over other labs and govs before getting noticed (other entities suspecting the lab getting close to ASI with non-neglible confidence) and reacting. This allows the lab to control the singleton instead of the mainstream labs and govs, and in this situation, the lab might escape from race dynamics.
However, this scenario results in a random lab controlling a singleton. While it’s not as hopeless as a singleton built by a racing entity, it doesn’t look really hopeful either.
If we’re talking about the differential effect of a given lab joining the race, then they could have a positive effect, if we know they have good intentions to benefit humanity. However, it’s still difficult to ensure the good intentions are still there when they actually get to ASI.
>If personas are a viable path to near-term alignment (and I think they are), control could set up a more adversarial relationship with the AI and increase the probability of misalignment that way.
I have some opinions on this (vibe based too):
if the AI is not very capable, then it could produce honest mistakes like https://incidentdatabase.ai/cite/1152/. Some control measures, like proper permission management, could help prevent this type of mistakes from escalating into disasters.
Even if control does more harm than good to overall safety of advanced AI, that doesn’t mean we should give it up now. We can use it now when it’s still beneficial, then drop it when it’s no longer so.
Overall I think this argument makes sense, and has some truth to it, but it doesn’t mean we should give up on control now.