I have a vague impression that others in the community think that deception in particular is much more central than I think it is, so I want to warn against that interpretation here: I think deception is an important problem, but its main importance is as an example of some broader issues in alignment.
I think deception is pretty central in that if we knew how to get an AI that was always honest and transparent: it always reported everything that it knows that we would consider relevant and important if we also knew it, we would have overcome the bulk of the danger.
Also, I agree that this “sub-problem” seems alignment complete, because the part of the AI that evaluates what the AI knows has to be able to, itself, understand everything the AI knows, and honestly report everything that’s relevant according to the standards of the human users. eg It seems like that part of the AI needs to itself be an aligned superintelligence.
Insofar as that’s true, carving out this “sub-problem” doesn’t help you at all.
But maybe it’s not true, and there’s some way of making an AI reliably honest and transparent that doesn’t route through having otherwise solved alignment?
I think deception is pretty central in that if we knew how to get an AI that was always honest and transparent: it always reported everything that it knows that we would consider relevant and important if we also knew it, we would have overcome the bulk of the danger.
Also, I agree that this “sub-problem” seems alignment complete, because the part of the AI that evaluates what the AI knows has to be able to, itself, understand everything the AI knows, and honestly report everything that’s relevant according to the standards of the human users. eg It seems like that part of the AI needs to itself be an aligned superintelligence.
Insofar as that’s true, carving out this “sub-problem” doesn’t help you at all.
But maybe it’s not true, and there’s some way of making an AI reliably honest and transparent that doesn’t route through having otherwise solved alignment?