Thanks, this is a great comment. I agree that your conclusions follow from your description of the situation, but your situation builds in a particular payment scheme: the AI gets paid if and only if we humans buy its alignment plan.
What we propose is a different payment scheme (see secs 4.2-6): the AI gets paid if and only if the AI does what we ask of it and isn’t caught doing anything suspicious. The AI gets this payment regardless of whether we buy its plan. The aim with that payment is to stop the AI rebelling. Then, for alignment work, the AI can also get a bonus later on, once we’ve reached a position where we can accurately evaluate its work. The aim with that bonus is to incentivize good work over slop.
If we go for that payment scheme, giving honest advice beats both rebelling and giving sycophantic advice. The outcomes of giving honest advice become X% good+bonus (if humans survive and later approve of the AI’s work), and 1-X% mitigated catastrophe (if humans don’t survive or don’t later approve). That makes giving honest advice dominate giving sycophantic advice: the AI gets the $500 no matter what, and if it gives honest advice there’s also an X% chance it gets the bonus on top of that. This is pretty incentive-compatible. The risk-averse AI wants to maximize its chances of getting the bonus, which means maximizing the chances that we humans survive and later approve of the AI’s work. This incentive-compatibility means we don’t have to lean so hard on our ability to recognize good alignment plans, because it’d let us trust that risk-averse AIs are truly trying their hardest (though of course better recognition of alignment-plan quality would still be very valuable).
And if (and only if) the AI is risk-averse, giving honest advice beats rebelling too, because the AI cares a huge amount about the difference between mitigated and unmitigated catastrophe, and cares very little about the difference between ideal and good+bonus.
Thanks, this is a great comment. I agree that your conclusions follow from your description of the situation, but your situation builds in a particular payment scheme: the AI gets paid if and only if we humans buy its alignment plan.
What we propose is a different payment scheme (see secs 4.2-6): the AI gets paid if and only if the AI does what we ask of it and isn’t caught doing anything suspicious. The AI gets this payment regardless of whether we buy its plan. The aim with that payment is to stop the AI rebelling. Then, for alignment work, the AI can also get a bonus later on, once we’ve reached a position where we can accurately evaluate its work. The aim with that bonus is to incentivize good work over slop.
If we go for that payment scheme, giving honest advice beats both rebelling and giving sycophantic advice. The outcomes of giving honest advice become X% good+bonus (if humans survive and later approve of the AI’s work), and 1-X% mitigated catastrophe (if humans don’t survive or don’t later approve). That makes giving honest advice dominate giving sycophantic advice: the AI gets the $500 no matter what, and if it gives honest advice there’s also an X% chance it gets the bonus on top of that. This is pretty incentive-compatible. The risk-averse AI wants to maximize its chances of getting the bonus, which means maximizing the chances that we humans survive and later approve of the AI’s work. This incentive-compatibility means we don’t have to lean so hard on our ability to recognize good alignment plans, because it’d let us trust that risk-averse AIs are truly trying their hardest (though of course better recognition of alignment-plan quality would still be very valuable).
And if (and only if) the AI is risk-averse, giving honest advice beats rebelling too, because the AI cares a huge amount about the difference between mitigated and unmitigated catastrophe, and cares very little about the difference between ideal and good+bonus.