cobylk
cobylk’s Shortform
I think this misplaces the left side of the window where control is useful. In particular, there will be a somewhat large period where AIs are insufficiently intelligent to successfully execute a takeover, but still could act on misaligned, coherent, long-term goals. Insofar as control approaches can in fact catch or mitigate those actions, there are lots of scenarios where control applied in the earlier stages of RSI is counterfactual for ensuring that later, takeover-capable models are more aligned and/or more poorly positioned to escape control. For example, control could prevent early schemers from sandbagging, sabotaging, or hijacking alignment research directions, setting up rogue deployments, and so on, before we get to takeover-capable AIs and the start of your control window.
Do we have reason to believe we will be able to reduce uncertainty about lab lead times? Currently, it seems unclear even to insiders which lab has the best internal model for AI R&D. This could be because the lead between the labs is very small (i.e., <1 month) or because we have very wide uncertainty about lab lead times. If it’s the latter and things don’t improve, this could pose an issue in many situations where we would like the lead lab would like to slow down and “burn” their lead to prioritize safety research (i.e., Plan C). See also The US-China model gap is very uncertain.