Largely disagree with this specific update, in large part because it’s very easy to increase p(survival) without p(flourishing), and the assumptions that make them strongly linked rely on premises that are relatively dubious, to put it mildly.
One of the larger takeaways I got from the Better Futures series as well as the Beyond Existential Risk post is that increasing survival probabilities do not increase flourishing probabilities by default, or at the very least that the connection between the probabilities is a lot weaker than often assumed, and from a flourishing perspective, increasing survival isn’t as useful as flourishing focused interventions.
(For one specific way this matters here, from a flourishing perspective, it’s very important that pacing the frontier does not become an effective indefinite pause of frontier AI, and that slowing down the intelligence explosion is better than pausing AI progress outright or continuing to accelerate progress relentlessly.)
I think embedded evaluators/transparency being mandated by law is already very good, but yes the pacing of the frontier in practice (assuming regulations don’t come which specifically ban or limit AI progress, which is now somewhat dubious) is probably going to be pretty light.
I expect somewhat more pacing than you do, and for your example, I’d say it’d be released in mid-to-late 2029 or 2030 instead of early 2029, especially in Anthropic, mostly because I trust the companies somewhat more than you do (especially Anthropic), but yes I agree this is the most likely way things end up post-Coxon incident, especially because I believe that the visible warning shots will decrease to 0, partially downstream of improved alignment, and partially because of improved eval awareness/metagaming, and also because the current incident seems likely to be because of specific and identifiable in-practice broken RL environments, and I think this likely generalizes somewhat well.