I think Plan A is significantly worse than Plan S.
I am uncertain that even the best versions of the sorts of control/superficial alignment techniques portrayed in the plan will be sufficient to make sure nothing catastrophic happens, when “Top-Expert-Dominating AI” is on the table.
And it does not seem to me that every time, across years and dozens of companies, the best possible versions of said alignment techniques will be implemented. It seems very plausible that something of the flavor of Anthropic and OpenAI training against the CoT will happen, except much more dangerous, since “Top-Expert-Dominating AI”.
I do not think that scaling to what Plan A portrays, will speed up the time to “alignment is solved” sufficiently to outweigh the risk; in my experience, the sorts of things current AIs are good at, or are on track to become good at, are not the limiting factor in alignment research, especially the sorts of exceptional alignment research that bring us substantially closer to “alignment is solved”.
I agree that deal breakdown is a big problem, one that people should be working on post-pause. I do not think getting closer to the edge of existentially dangerous capabilities helps with that.[1]
- ^
If anything, it might make a deal weaker, since it goes past a natural Schelling fence.
Not enough though, I think. “smart people are attracted into unproductive directions because they stick to what’s memetic, avoid conflicting with allies, and are too panicked to think carefully” does describe many people in the EA-sphere. More is possible.