I think embedded evaluators/transparency being mandated by law is already very good, but yes the pacing of the frontier in practice (assuming regulations don’t come which specifically ban or limit AI progress, which is now somewhat dubious) is probably going to be pretty light.
I expect somewhat more pacing than you do, and for your example, I’d say it’d be released in mid-to-late 2029 or 2030 instead of early 2029, especially in Anthropic, mostly because I trust the companies somewhat more than you do (especially Anthropic), but yes I agree this is the most likely way things end up post-Coxon incident, especially because I believe that the visible warning shots will decrease to 0, partially downstream of improved alignment, and partially because of improved eval awareness/metagaming, and also because the current incident seems likely to be because of specific and identifiable in-practice broken RL environments, and I think this likely generalizes somewhat well.
My current take here is that for the purposes of succession, you actually don’t need to solve moral philosophy in full, forever, mostly because of the possibility of moral trade/compromise both causally and acausally that lets most people get most of what they want while only meeting the first bar.
(To be clear, I’m not a successionist in the sense that I don’t think it’s acceptable to intentionally increase the probability we all die because you think human extinction is good/fine, but of course all of the debates is between non-successionists in the way I defined it.)
While I don’t think liberalism or democracy will survive AI progress, I do think that it’s (maybe) tractable, and important/neglected to port over one of the concepts/founding reasons for liberalism and democracy, which is attempting to compromise/trade across people with vastly different moral views.
Or in the words of John Rawls:
And the answer for the future is trade/compromise can replace voting/electoral systems.
This still requires a lot of improvement in the political situation, such that people can consider trade/compromise to be worth it rather than their easy reaction of banning/trying to inflict violence on opposing groups, but contra your essay, I don’t believe the succession problem is as hard as you state.
(Some other reasons why the succession problem is easy, conditional on having corrigible AIs is because unlike humans, who grow old and die, corrigible AIs will be living for a very, very long time, and their only limit is how much free energy they can get/generate, and once they die, the universe is maximally boring and in thermodynamic equlibrium, where no more work is possible, or the decay of the vacuum happened and killed everyone, and not being able to live forever is part of my explanation for why dictatorships eventually trend towards the same bad endpoint from a citizen perspective, and relatedly it’s much easier to lock-in desirable values post-AGI.)