If the goal is to hand off to AI early-ish (which I’m not claiming is good or bad) then even some prosaic alignment things like trying to systematically understand midtraining or trying to “solve” eval awareness seem less risky while being equally productive and maybe more tractable.
I think there are even more prosaic things we could could do to make handoff go well!
In rough order of confidence that the input will be relevant, automated safety work requires:
Access to compute;
Access to frontier or near-frontier models; and of course,
Aligned, controlled, and capable models.
Compute and model access strike me as P0 for handoff, necessary in the biggest % of takeoff worlds, as well as carrying fewer differential downside risks.
I think there are even more prosaic things we could could do to make handoff go well!
In rough order of confidence that the input will be relevant, automated safety work requires:
Access to compute;
Access to frontier or near-frontier models; and of course,
Aligned, controlled, and capable models.
Compute and model access strike me as P0 for handoff, necessary in the biggest % of takeoff worlds, as well as carrying fewer differential downside risks.
Very good point! I didn’t consider this framing.