Harry Waterman
Karma: 21
Concrete Generalist Projects in AI Safety (and how to do them)
Harry Waterman’s Shortform
If the goal is to hand off to AI early-ish (which I’m not claiming is good or bad) then even some prosaic alignment things like trying to systematically understand midtraining or trying to “solve” eval awareness seem less risky while being equally productive and maybe more tractable.
I think there are even more prosaic things we could could do to make handoff go well!
In rough order of confidence that the input will be relevant, automated safety work requires:
Access to compute;
Access to frontier or near-frontier models; and of course,
Aligned, controlled, and capable models.
Compute and model access strike me as P0 for handoff, necessary in the biggest % of takeoff worlds, as well as carrying fewer differential downside risks.
We should expect ASI to indeed be radically transformative, to the point where our usual human abstractions like “multipolarity” or “authoritarianism” have very little meaning. This intuition caused me to dismiss most power-concentration threat models as sort of confused, or at least overconfident in their descriptions of future dynamics.
One objection, though, is that power-seeking humans may have an incentive to restrain the transformative properties of AI at a level that preserves human-shaped things like polarity or authoritarianism, given they benefit. So we shouldn’t be too surprised if we end up in a hellish power-concentrated world (absent loss of control).