If failure of alignment → schemer that wants to seize power, then ASI alignment is one shot.
But if failure of alignment → non-schemer misalignment (eg reward hacking, or flailing misgeneralisation), then we failure isn’t existential
So I think p(scheming | alignment failure) is a crux here
If failure of alignment → schemer that wants to seize power, then ASI alignment is one shot.
But if failure of alignment → non-schemer misalignment (eg reward hacking, or flailing misgeneralisation), then we failure isn’t existential
So I think p(scheming | alignment failure) is a crux here