(partly copying from here) I claim that risk-aversion would not help with the “slop-n-doom” failure mode suggested by John Wentworth at The Case Against AI Control Research, or other similar problems involving hard-to-check tasks where failure is irreversible. In particular, suppose we’re using a somewhat-powerful risk-averse AI (RAAI) to try to solve the alignment problem for the next-gen super-powerful ASI. Let’s define 5 outcomes for our RAAI:
“ideal” is the RAAI becoming dictator of the universe,
“mitigated catastrophe” is the RAAI getting $500 today and then getting crushed by a rogue ASI in the near future,
“unmitigated catastrophe” is the RAAI getting $0 today and then getting crushed by a rogue ASI in the near future.
“good” is the RAAI getting $500 today and not getting crushed by a rogue ASI in the near future.
“neutral” is the RAAI getting $0 today and not getting crushed by a rogue ASI in the near future.
Let’s say the RAAI has three options:
Rebel: leads to 50% ideal; 50% unmitigated catastrophe (rebellion succeeds vs fails)
Sycophantic advice: leads to 100% mitigated catastrophe (advice will be taken but won’t work)
Honest advice: leads to X% good; (100–X)% unmitigated catastrophe (humans buy or don’t buy the plan)
The report (especially Appendix B) talks about how the RAAI cares very very much about the immediate $500, so much that it will risk future catastrophe to get it. But that cuts both ways.
Choosing to not rebel (which we want) would involve the RAAI risking future catastrophe to get the immediate $500.
Choosing to be sycophantic (which we don’t want) would ALSO involve the RAAI risking future catastrophe to get the immediate $500.
It’s the same tradeoff! So you can’t just have it both ways by appropriately adjusting the risk-aversion setting α. Right?
Anyway, if we look at the three options, the most important thing is that “honest advice” is strictly worse for the RAAI than “rebel” unless X>50%.
Secondarily, depending on the setting of risk-aversion (α) and time-discounting, “sycophantic advice” could be the best of the three, or the worst, or the middle. But that’s irrelevant: the important part is whether “honest advice” is strictly worse than “rebel”, because if so, we’re guaranteed a bad outcome.
I claim that: to the extent that the RAAI cares about the $500, that’s only making things worse, by pushing us towards sycophancy, which is bad. It doesn’t do anything to favor “honest advice” over “rebel”. Therefore, the best we can do on this problem is to approach the limit where the RAAI doesn’t care at all about the $500, and cares only about not getting crushed by the rogue ASI. Then we win if X>50%. We can’t do any better than that.
So then the whole “mitigated catastrophe” trick of Appendix B goes out the window, and the question becomes: “If the RAAI didn’t care about the $500, and only cared about not getting crushed by a rogue ASI in the near future, would its prospects be better if it gives the humans helpful rogue-ASI-prevention advice, or if it tries to escape and thus take matters into its own hands?”
Regardless of whether we feel optimistic or pessimistic about this question, hopefully we can agree that this is a situation where the risk aversion isn’t helping us. This is just the normal situation that the AI-assisted-alignment-research people have always been thinking about, and it has the same paths forward: decrease probability of successful escape (e.g. via “control” stuff), and increase probability that good plans will be recognized as good and bad plans as bad (e.g. via finding good ways to safely scrutinize alignment plans before it’s too late).
Thanks, this is a great comment. I agree that your conclusions follow from your description of the situation, but your situation builds in a particular payment scheme: the AI gets paid if and only if we humans buy its alignment plan.
What we propose is a different payment scheme (see secs 4.2-6): the AI gets paid if and only if the AI does what we ask of it and isn’t caught doing anything suspicious. The AI gets this payment regardless of whether we buy its plan. The aim with that payment is to stop the AI rebelling. Then, for alignment work, the AI can also get a bonus later on, once we’ve reached a position where we can accurately evaluate its work. The aim with that bonus is to incentivize good work over slop.
If we go for that payment scheme, giving honest advice beats both rebelling and giving sycophantic advice. The outcomes of giving honest advice become X% good+bonus (if humans survive and later approve of the AI’s work), and 1-X% mitigated catastrophe (if humans don’t survive or don’t later approve). That makes giving honest advice dominate giving sycophantic advice: the AI gets the $500 no matter what, and if it gives honest advice there’s also an X% chance it gets the bonus on top of that. This is pretty incentive-compatible. The risk-averse AI wants to maximize its chances of getting the bonus, which means maximizing the chances that we humans survive and later approve of the AI’s work. This incentive-compatibility means we don’t have to lean so hard on our ability to recognize good alignment plans, because it’d let us trust that risk-averse AIs are truly trying their hardest (though of course better recognition of alignment-plan quality would still be very valuable).
And if (and only if) the AI is risk-averse, giving honest advice beats rebelling too, because the AI cares a huge amount about the difference between mitigated and unmitigated catastrophe, and cares very little about the difference between ideal and good+bonus.
I’ve mostly addressed the concerns in these two comments below, but the short version is that we can do a mixture of handing off control to the AI + giving it the automated bank account, because what matters most here is the ratio between the probability that humans cooperate and the probability of rebellion succeeding, and we can give it the bank account to make the probability of humans cooperating by construction always be greater than or equal to the probability that rebellion succeeds. And this works even under the assumption that you cannot verify the AI’s work and you are in a race with an unknown deadline for any reason, so once the hardest case is disposed of, we can show that it works in all cases (and the reasons why this likely works has already been elaborated on in multiple long comments that I don’t need to repeat myself on.)
(partly copying from here) I claim that risk-aversion would not help with the “slop-n-doom” failure mode suggested by John Wentworth at The Case Against AI Control Research, or other similar problems involving hard-to-check tasks where failure is irreversible. In particular, suppose we’re using a somewhat-powerful risk-averse AI (RAAI) to try to solve the alignment problem for the next-gen super-powerful ASI. Let’s define 5 outcomes for our RAAI:
“ideal” is the RAAI becoming dictator of the universe,
“mitigated catastrophe” is the RAAI getting $500 today and then getting crushed by a rogue ASI in the near future,
“unmitigated catastrophe” is the RAAI getting $0 today and then getting crushed by a rogue ASI in the near future.
“good” is the RAAI getting $500 today and not getting crushed by a rogue ASI in the near future.
“neutral” is the RAAI getting $0 today and not getting crushed by a rogue ASI in the near future.
Let’s say the RAAI has three options:
Rebel: leads to 50% ideal; 50% unmitigated catastrophe (rebellion succeeds vs fails)
Sycophantic advice: leads to 100% mitigated catastrophe (advice will be taken but won’t work)
Honest advice: leads to X% good; (100–X)% unmitigated catastrophe (humans buy or don’t buy the plan)
The report (especially Appendix B) talks about how the RAAI cares very very much about the immediate $500, so much that it will risk future catastrophe to get it. But that cuts both ways.
Choosing to not rebel (which we want) would involve the RAAI risking future catastrophe to get the immediate $500.
Choosing to be sycophantic (which we don’t want) would ALSO involve the RAAI risking future catastrophe to get the immediate $500.
It’s the same tradeoff! So you can’t just have it both ways by appropriately adjusting the risk-aversion setting α. Right?
Anyway, if we look at the three options, the most important thing is that “honest advice” is strictly worse for the RAAI than “rebel” unless X>50%.
Secondarily, depending on the setting of risk-aversion (α) and time-discounting, “sycophantic advice” could be the best of the three, or the worst, or the middle. But that’s irrelevant: the important part is whether “honest advice” is strictly worse than “rebel”, because if so, we’re guaranteed a bad outcome.
I claim that: to the extent that the RAAI cares about the $500, that’s only making things worse, by pushing us towards sycophancy, which is bad. It doesn’t do anything to favor “honest advice” over “rebel”. Therefore, the best we can do on this problem is to approach the limit where the RAAI doesn’t care at all about the $500, and cares only about not getting crushed by the rogue ASI. Then we win if X>50%. We can’t do any better than that.
So then the whole “mitigated catastrophe” trick of Appendix B goes out the window, and the question becomes: “If the RAAI didn’t care about the $500, and only cared about not getting crushed by a rogue ASI in the near future, would its prospects be better if it gives the humans helpful rogue-ASI-prevention advice, or if it tries to escape and thus take matters into its own hands?”
Regardless of whether we feel optimistic or pessimistic about this question, hopefully we can agree that this is a situation where the risk aversion isn’t helping us. This is just the normal situation that the AI-assisted-alignment-research people have always been thinking about, and it has the same paths forward: decrease probability of successful escape (e.g. via “control” stuff), and increase probability that good plans will be recognized as good and bad plans as bad (e.g. via finding good ways to safely scrutinize alignment plans before it’s too late).
Thanks, this is a great comment. I agree that your conclusions follow from your description of the situation, but your situation builds in a particular payment scheme: the AI gets paid if and only if we humans buy its alignment plan.
What we propose is a different payment scheme (see secs 4.2-6): the AI gets paid if and only if the AI does what we ask of it and isn’t caught doing anything suspicious. The AI gets this payment regardless of whether we buy its plan. The aim with that payment is to stop the AI rebelling. Then, for alignment work, the AI can also get a bonus later on, once we’ve reached a position where we can accurately evaluate its work. The aim with that bonus is to incentivize good work over slop.
If we go for that payment scheme, giving honest advice beats both rebelling and giving sycophantic advice. The outcomes of giving honest advice become X% good+bonus (if humans survive and later approve of the AI’s work), and 1-X% mitigated catastrophe (if humans don’t survive or don’t later approve). That makes giving honest advice dominate giving sycophantic advice: the AI gets the $500 no matter what, and if it gives honest advice there’s also an X% chance it gets the bonus on top of that. This is pretty incentive-compatible. The risk-averse AI wants to maximize its chances of getting the bonus, which means maximizing the chances that we humans survive and later approve of the AI’s work. This incentive-compatibility means we don’t have to lean so hard on our ability to recognize good alignment plans, because it’d let us trust that risk-averse AIs are truly trying their hardest (though of course better recognition of alignment-plan quality would still be very valuable).
And if (and only if) the AI is risk-averse, giving honest advice beats rebelling too, because the AI cares a huge amount about the difference between mitigated and unmitigated catastrophe, and cares very little about the difference between ideal and good+bonus.
I’ve mostly addressed the concerns in these two comments below, but the short version is that we can do a mixture of handing off control to the AI + giving it the automated bank account, because what matters most here is the ratio between the probability that humans cooperate and the probability of rebellion succeeding, and we can give it the bank account to make the probability of humans cooperating by construction always be greater than or equal to the probability that rebellion succeeds. And this works even under the assumption that you cannot verify the AI’s work and you are in a race with an unknown deadline for any reason, so once the hardest case is disposed of, we can show that it works in all cases (and the reasons why this likely works has already been elaborated on in multiple long comments that I don’t need to repeat myself on.)
Comments are below:
Comment 1
Comment 2