These predictions are not very related to any alignment research that is currently occurring. I think it’s just quite unclear how hard the problem is, e.g. does deceptive alignment occur, do models trained to honestly answer easy questions generalize to hard questions, how much intellectual work are AI systems doing before they can take over, etc.
I know people have spilled a lot of ink over this, but right now I don’t have much sympathy for confidence that the risk will be real and hard to fix (just as I don’t have much sympathy for confidence that the problem isn’t real or will be easy to fix).
These predictions are not very related to any alignment research that is currently occurring. I think it’s just quite unclear how hard the problem is, e.g. does deceptive alignment occur, do models trained to honestly answer easy questions generalize to hard questions, how much intellectual work are AI systems doing before they can take over, etc.
I know people have spilled a lot of ink over this, but right now I don’t have much sympathy for confidence that the risk will be real and hard to fix (just as I don’t have much sympathy for confidence that the problem isn’t real or will be easy to fix).