In worlds where AI alignment can be handled by iterative design, we probably survive.
I’m curious whether you still believe in this? I think there are currently lots of alignment issues that are totally fixable by iterating on them, but they aren’t. Naively extrapolating, it’s possible that even if various future alignment problems are solvable via iteration, they might not be.
My guess is that at the current state, conditional on (misalignment at the point of no return will kill everyone) AND (this misalignment is fixable using iteration), p(misalignment) is still around 50%.[1]
Obvious this is a bit fuzzy since iterative design worlds would have much more continuous looking PONR and misalignment issues, but that’s the sort of vibe I have about the current state of the AI race.
I’m curious whether you still believe in this? I think there are currently lots of alignment issues that are totally fixable by iterating on them, but they aren’t. Naively extrapolating, it’s possible that even if various future alignment problems are solvable via iteration, they might not be.
My guess is that at the current state, conditional on (misalignment at the point of no return will kill everyone) AND (this misalignment is fixable using iteration), p(misalignment) is still around 50%.[1]
Obvious this is a bit fuzzy since iterative design worlds would have much more continuous looking PONR and misalignment issues, but that’s the sort of vibe I have about the current state of the AI race.