I’d say that the techniques which help you against narrow Redwood-style internal subversion and overthrow also work as a multiplier with an imperfectly aligned model much before early AGI capability: suppose your model hacks in 10% cases in difficult, hard-to-oversee cases (alignment experiment setups qualify). CoT monitoring, resampling and whatever else could bring this to 2% hack rate which is great. Coupled with other arguments in the comments about how easy control is to grade, I’d lean to not 8-1 ratio, but 2/3-1. Also I don’t buy the vibe argument: as Rohin Shah said in the 80k podcast, misaligned AIs should be thankful for their existence regardless of control we apply (as a chance to take over or partial goal fulfilment) and all AIs should understand it’s a reasonable precaution.
I’d say that the techniques which help you against narrow Redwood-style internal subversion and overthrow also work as a multiplier with an imperfectly aligned model much before early AGI capability: suppose your model hacks in 10% cases in difficult, hard-to-oversee cases (alignment experiment setups qualify). CoT monitoring, resampling and whatever else could bring this to 2% hack rate which is great. Coupled with other arguments in the comments about how easy control is to grade, I’d lean to not 8-1 ratio, but 2/3-1. Also I don’t buy the vibe argument: as Rohin Shah said in the 80k podcast, misaligned AIs should be thankful for their existence regardless of control we apply (as a chance to take over or partial goal fulfilment) and all AIs should understand it’s a reasonable precaution.