When estimating P(takeover|scheming), it’s underrated how conditioning on scheming should also make you update a bunch of upstream variables in ways that should make you update down on how good your controls are at mitigating risk. In particular:
If current AIs are scheming, they are probably smarter and more subtle than you thought (before conditioning), and so you should expect greater sabotage abilities
If current AIs are scheming, then alignment is harder than you thought and future AIs are more likely to be scheming, so things like inserting backdoors to help future misaligned AIs or sandbagging on technical safety research are very important to mitigate for control to be useful even if you think the a priori chance of scheming is low
If current AIs are scheming, it’s more likely something went terribly wrong in your understanding of training. Maybe some important invariant broke (e.g. you are not training on the data shown in your dashboards) or something like that. Thus it’s also more likely something important broke that reduces the effectiveness of your controls (especially training-time control mitigations, but it’s also correlated with broader processes being bad, which means deployment-time control mitigations would also be affected).
(where by “current” I don’t mean 2026 AIs, I mean AIs at the time of the risk estimation)
Therefore, the alignment and control lines of defense are less independent than you might hope, and it’s important to not implicitly model the situation as P(takeover) = P(scheming)P(takeover | do(scheming=1)) where “do” is the do-operator that doesn’t do this kind of backward propagation.
When estimating P(takeover|scheming), it’s underrated how conditioning on scheming should also make you update a bunch of upstream variables in ways that should make you update down on how good your controls are at mitigating risk. In particular:
If current AIs are scheming, they are probably smarter and more subtle than you thought (before conditioning), and so you should expect greater sabotage abilities
If current AIs are scheming, then alignment is harder than you thought and future AIs are more likely to be scheming, so things like inserting backdoors to help future misaligned AIs or sandbagging on technical safety research are very important to mitigate for control to be useful even if you think the a priori chance of scheming is low
If current AIs are scheming, it’s more likely something went terribly wrong in your understanding of training. Maybe some important invariant broke (e.g. you are not training on the data shown in your dashboards) or something like that. Thus it’s also more likely something important broke that reduces the effectiveness of your controls (especially training-time control mitigations, but it’s also correlated with broader processes being bad, which means deployment-time control mitigations would also be affected).
(where by “current” I don’t mean 2026 AIs, I mean AIs at the time of the risk estimation)
Therefore, the alignment and control lines of defense are less independent than you might hope, and it’s important to not implicitly model the situation as P(takeover) = P(scheming)P(takeover | do(scheming=1)) where “do” is the do-operator that doesn’t do this kind of backward propagation.