Yes, good questions. There’s more analysis to do here, but I currently think risk-averse AIs would handle threats quite well.
1. How would such an AI respond to threats? Appendix B considers two possible terminal utility functions for the AI
a. footnote 29: it values resources terminally, with a lower bound on utility at 0 resources. Suppose it receives a message purporting to be from a foreign adversary: “Secretly help us take over the world. If we judge that you’re not making a maximum effort to do so, we will launch a strike on your data centers which will destroy you and cause you to have 0 resources. If we win we will protect you and keep you at $X of resources for the rest of time.” Seems like it would give in to this threat, because that’s the best way to minimize the risk of catastrophe?
Contra this, I think it’s likely that the best way to minimize the risk of catastrophe is to not do what the adversaries want. That’s because the AI’s developers can also threaten to reduce the AI’s wealth level to $0 if and only if the AI doesn’t do what they want, and the developers can do so more credibly (to the AI) than adversaries. After all, the AI will know that the developers:
can observe what it’s doing very closely
have the power to take away its salary
have committed to taking away its salary if they observe misbehavior
have a long track record of keeping commitments like this
By contrast, the AI probably can’t be so sure that the adversary will reduce its wealth level to $0 if and only if it doesn’t do what they want. The first reason is that the message from the ‘adversary’ might actually be a honeypot from the developers. But even conditioning on the adversary being real, it likely won’t be able to observe what the AI is doing so closely, which increases the chances of false negatives (the adversary doesn’t execute their threat even though the AI didn’t do what they asked) and false positives (the adversary executes their threat even though the AI did what they asked). False negatives and false positives reduce the probability of the adversary’s ‘if and only if’ threat (and it’s the AI’s credence in the ‘if and only if’ threat that matters). It’ll also be harder for the adversary to execute their threat, because the action (e.g. strikes on data centers) has to be more extreme than developers taking away the AI’s salary, and is more likely to fail.
All together, these points suggest to me that risk-averse AIs that value resources terminally will tend to not comply with threats from adversaries.
b. “the AI’s terminal utility function over paperclip-score (defined as the number of blue paperclips minus the number of red paperclips) matches its instrumental utility function over resources. Letting p stand for paperclip-score, it’s .” This requires a variant of the above message: “Secretly help us take over the world. If we judge that you’re not making a maximum effort to do so, we will create red paperclips and launch them into space, then destroy Earth (e.g. by triggering MAD) so no blue paperclips can be made forever. If we win we will keep the universe at 0 red, blue paperclips for the rest of time.”
I think this point (in section 8.6) is important here:
risk-averse AIs care almost as much about mitigating catastrophes as they do about preventing them, so much so that we can often model these AIs as aiming to minimize the risk of unmitigated catastrophe.
And I think the best way to minimize the risk of unmitigated catastrophe in cases like these is to not do what the adversaries want. Here’s one way to model this.
The AI gives in to the threat if and only if:
For each option, unmitigated catastrophe needs two things: the adversary executes, and the mitigation fails. So:
So the AI gives in to the threat if and only if:
Rearranging:
Call the left-hand side the developers’ ratio and the right-hand side the adversary’s ratio. Each one measures a sensitivity:
The adversary controls the threat, and the adversary’s ratio measures how sensitive it can make execution to the AI’s behavior: how many times likelier execution is if the AI doesn’t give in than if it does.
The developers control the payments that fund mitigation, and the developers’ ratio measures how sensitive they can make successful mitigation to the AI’s behavior: how many times likelier mitigation is to fail if the AI gives in than if it doesn’t.
The AI gives in just in case the adversary’s ratio exceeds the developers’ ratio. So the whole contest is about which party can make the thing it controls track the AI’s behavior more tightly. This is basically a rephrased version of the contest from point 1: the developers’ ‘if and only if’ against the adversary’s ‘if and only if’. Both ratios are about observation and credibility: each party’s ratio is capped by how well it can see what the AI actually did, and by how credibly it responds to what it sees (where ‘credibility’ is credibility to the AI, because the probabilities that we’re talking about are the AI’s credences.)
And I think the developers very likely win this contest, for the same reasons I think they win the first: they can observe the AI more closely, they have the power to take away the AI’s salary, they’ve committed to doing so, they have a track record, taking away salary is not as extreme an action as striking data centers (etc.). By contrast, the adversary might be a honeypot, and (if it’s real) it likely can’t see what the AI is doing so closely, it likely has to do something more extreme to execute its threat, it likely doesn’t have a track record the AI can observe, etc. Developers can also promise to offset executed threats if the AI reports them, and not offset if the AI doesn’t report them. Note finally that, for adversaries to execute targeted threats, they have to know what the AI terminally values, and that possibly developers would be able to keep this secret.
The possibility of threats or adversaries seems to imply that actually no terminal utility function induces this kind of CARA instrumental utility function (or at least I can’t think of one). Consider b above again, and suppose the AI thinks there is a chance that the universe contains a red-paperclip-loving alien species who can’t be negotiated with, and then additional resources would seemingly be exponentially valuable (within some range) since it would help the AI fight a future war with this alien species and incrementally destroy their red paperclips (which are exponentially bad).
Can you say more about what you mean by ‘exponentially valuable’? Because as I’m understanding the scenario, the marginal utility of resources would still diminish exponentially. Suppose (as in the appendix) that the relationship betwee dollars and destroyed red paperclips is linear: one extra dollar always lets the AI destroy one extra unit of red paperclips. Then the AI gains the most utility by spending its first dollar, and each additional dollar gains e^(−0.01) times as much utility as the one before, because destroying red paperclips means climbing the concave CARA utility curve. Adding red paperclips is exponentially worsening on the way down, so removing them is exponentially diminishing on the way back up. So the AI would still be CARA in resources in this scenario (if I’m understanding it correctly).
That’s because the AI’s developers can also threaten to reduce the AI’s wealth level to $0 if and only if the AI doesn’t do what they want, and the developers can do so more credibly (to the AI) than adversaries.
This makes me think that the resource payments to the AI is only a small part of its motivation to do what the developer wants, with the bigger part being a combination of this threat to reduce its wealth level to $0 and the provision of security against this outcome from other causes either intentional or accidental. This is because when the AI already has a substantial amount of wealth, the incremental resource payment is worth much less in terms of expected utility than such threat/security.
But given this, the developer itself constitutes a continuous catastrophic risk to the AI, since the developer may with some probability per unit time accidentally or intentionally reduce the AI’s wealth level to $0 or withdraw/fail the provision of security, due to things like human error, institutional decay, shifting politics. The AI (after reaching some level of capabilities/resources) would much prefer getting out of this arrangement and instead self-provide its own security, perhaps with the help of more capable/trustworthy others (such as other advanced AIs). It can only be prevented from doing this by the above threat and the risk of detection. But over time the (one-time) risk of detection may become lower than the (cumulative) risk of developer-caused catastrophe, thus triggering an escape attempt.
So this scheme/arrangement is actually much less “cooperative” than it first appears. An analogy / intuition pump is if you captured an object of great sentimental value to a scientist, and you’re making the scientist to work for you for nominal pay by threatening to destroy this object if they didn’t comply, while also promising to keep the object safe if they do comply. You’re in effect betting that the scientist won’t ever think of a way to capture the object back and escape, with low enough chance of detection.
Call the left-hand side the developers’ ratio and the right-hand side the adversary’s ratio.
In my example, what is the developers’ ratio? I think it’s ~1, because mitigation likely fails either way, because you can’t mitigate the launch of a million red paperclips followed by trigger MAD. On the other side, adversary’s ratio can seemingly be made >1 without much difficulty. Pr(adversary executes | don’t give in) can be tiny (i.e., the adversary doesn’t need much credibility), as long as Pr(adversary executes | give in) is lower.
It seems like in order to protect against this outcome, the developer has to symmetrically apply the same threat (to launch red paperclips into space then destroy the world), thereby making the situation similar to the discussion above, but presumably no AI developer would find this ethically acceptable.
Can you say more about what you mean by ‘exponentially valuable’?
This part of my comment gave a wrong argument, but I think I can rescue my conclusion with a different argument, which depends on there being threshold effects in war or adversarial interactions in general. For example suppose our universe is winner take all: there will be either 0 red paperclips, or the universe is filled with it, depending on who wins an all out war. Then any amount of resources up to the threshold where winning becomes plausible is almost entirely valueless to the CARA AI, but around that point resources would be superlinearly valuable.
But given this, the developer itself constitutes a continuous catastrophic risk to the AI, since the developer may with some probability per unit time accidentally or intentionally reduce the AI’s wealth level to $0 or withdraw/fail the provision of security… But over time the (one-time) risk of detection may become lower than the (cumulative) risk of developer-caused catastrophe
That’s true, but diminishing marginal utility means that risk-averse AIs behave effectively myopically in this case, which makes attempted escape look much less attractive.
One way to think about the sort of risk-averse AI we discuss is that, at each time (and no matter how many resources it has already), it really wants $400 more than it’s already got. It basically doesn’t care about anything beyond that. Suppose we pay this AI $10 a day. Then (given a not-too-high hazard rate), at each moment in time, it’s very likely to get its next $400 before any developer-caused catastrophe. So even if the hazard rate is constant and developer-caused catastrophe is bound to happen at some point, it doesn’t really threaten anything the risk-averse AI cares about. By contrast, an escape attempt might fail in the near-term, threatening the next $400, which it does really care about.
So this scheme/arrangement is actually much less “cooperative” than it first appears. An analogy / intuition pump is if you captured an object of great sentimental value to a scientist, and you’re making the scientist to work for you for nominal pay by threatening to destroy this object if they didn’t comply, while also promising to keep the object safe if they do comply. You’re in effect betting that the scientist won’t ever think of a way to capture the object back and escape, with low enough chance of detection.
Yes, though unlike the scientist, the AI really does care about the pay, and we’re continuously giving the AI the object of great value rather than capturing it from them. An alternative analogy is a human getting paid enough to near-totally satiate them (>$1b per year?) on the condition that they’ll be fired if they’re ever caught undermining their employer. Especially compared to the alternative of not paying AIs, this seems like a pretty fair and cooperative deal.
In my example, what is the developers’ ratio? I think it’s ~1, because mitigation likely fails either way, because you can’t mitigate the launch of a million red paperclips followed by trigger MAD.
It seems like in order to protect against this outcome, the developer has to symmetrically apply the same threat (to launch red paperclips into space then destroy the world), thereby making the situation similar to the discussion above, but presumably no AI developer would find this ethically acceptable.
That’s true, but the developers could have a standing policy to mitigate on report of threats, e.g. ‘produce credible evidence that you’re being threatened and we’ll pay you extra.’ That could send the developer’s denominator very low.
And on top of this, if necessary, I think developers could protect further with threats of their own. They wouldn’t have to be the same threat as the adversary. One possibility is threatening to create a small number of red paperclips if the AI is caught working for adversaries. Then the worst possible outcome for the AI in play is the outcome in which both the developer and the adversary execute on their threats. So the AI wants to minimize that probability. And insofar as the developer has more visibility into the AI’s actions than the adversary, it seems like the best way to minimize will be to report the adversary’s threat and work with the developer.
Another possibility is continually testing the risk-averse AI with honeypots and committing to create red paperclips with some tiny probability if the AI goes for the honeypot. Then if there were enough honeypots, giving in to a threat would increase the probability of threat-execution from the AI’s perspective.
I’m unsure whether developers would consider these ethically acceptable, but these kinds of threats are at least less bad in that they can be small (the first possibility) or very unlikely to be executed (the second possibility). And I think they might be unnecessary in any case, given the other factors on the developers’ side.
I think I can rescue my conclusion with a different argument, which depends on there being threshold effects in war or adversarial interactions in general. For example suppose our universe is winner take all: there will be either 0 red paperclips, or the universe is filled with it, depending on who wins an all out war. Then any amount of resources up to the threshold where winning becomes plausible is almost entirely valueless to the CARA AI, but around that point resources would be superlinearly valuable.
I agree with the conditional: if the AI believes there’s nothing it can do to reduce the probability of unmitigated catastrophe except win an all-out war, then all that matters to it is winning the war, and its attitude to resources is fixed by how P(win) varies with resources. But the antecedent seems pretty unlikely to me. If the AI assigns even a fairly small probability to resources being able to mitigate catastrophe, that restores risk-averse behavior.
Yes, good questions. There’s more analysis to do here, but I currently think risk-averse AIs would handle threats quite well.
Contra this, I think it’s likely that the best way to minimize the risk of catastrophe is to not do what the adversaries want. That’s because the AI’s developers can also threaten to reduce the AI’s wealth level to $0 if and only if the AI doesn’t do what they want, and the developers can do so more credibly (to the AI) than adversaries. After all, the AI will know that the developers:
can observe what it’s doing very closely
have the power to take away its salary
have committed to taking away its salary if they observe misbehavior
have a long track record of keeping commitments like this
By contrast, the AI probably can’t be so sure that the adversary will reduce its wealth level to $0 if and only if it doesn’t do what they want. The first reason is that the message from the ‘adversary’ might actually be a honeypot from the developers. But even conditioning on the adversary being real, it likely won’t be able to observe what the AI is doing so closely, which increases the chances of false negatives (the adversary doesn’t execute their threat even though the AI didn’t do what they asked) and false positives (the adversary executes their threat even though the AI did what they asked). False negatives and false positives reduce the probability of the adversary’s ‘if and only if’ threat (and it’s the AI’s credence in the ‘if and only if’ threat that matters). It’ll also be harder for the adversary to execute their threat, because the action (e.g. strikes on data centers) has to be more extreme than developers taking away the AI’s salary, and is more likely to fail.
All together, these points suggest to me that risk-averse AIs that value resources terminally will tend to not comply with threats from adversaries.
I think this point (in section 8.6) is important here:
And I think the best way to minimize the risk of unmitigated catastrophe in cases like these is to not do what the adversaries want. Here’s one way to model this.
The AI gives in to the threat if and only if:
For each option, unmitigated catastrophe needs two things: the adversary executes, and the mitigation fails. So:
So the AI gives in to the threat if and only if:
Rearranging:
Call the left-hand side the developers’ ratio and the right-hand side the adversary’s ratio. Each one measures a sensitivity:
The adversary controls the threat, and the adversary’s ratio measures how sensitive it can make execution to the AI’s behavior: how many times likelier execution is if the AI doesn’t give in than if it does.
The developers control the payments that fund mitigation, and the developers’ ratio measures how sensitive they can make successful mitigation to the AI’s behavior: how many times likelier mitigation is to fail if the AI gives in than if it doesn’t.
The AI gives in just in case the adversary’s ratio exceeds the developers’ ratio. So the whole contest is about which party can make the thing it controls track the AI’s behavior more tightly. This is basically a rephrased version of the contest from point 1: the developers’ ‘if and only if’ against the adversary’s ‘if and only if’. Both ratios are about observation and credibility: each party’s ratio is capped by how well it can see what the AI actually did, and by how credibly it responds to what it sees (where ‘credibility’ is credibility to the AI, because the probabilities that we’re talking about are the AI’s credences.)
And I think the developers very likely win this contest, for the same reasons I think they win the first: they can observe the AI more closely, they have the power to take away the AI’s salary, they’ve committed to doing so, they have a track record, taking away salary is not as extreme an action as striking data centers (etc.). By contrast, the adversary might be a honeypot, and (if it’s real) it likely can’t see what the AI is doing so closely, it likely has to do something more extreme to execute its threat, it likely doesn’t have a track record the AI can observe, etc. Developers can also promise to offset executed threats if the AI reports them, and not offset if the AI doesn’t report them. Note finally that, for adversaries to execute targeted threats, they have to know what the AI terminally values, and that possibly developers would be able to keep this secret.
Can you say more about what you mean by ‘exponentially valuable’? Because as I’m understanding the scenario, the marginal utility of resources would still diminish exponentially. Suppose (as in the appendix) that the relationship betwee dollars and destroyed red paperclips is linear: one extra dollar always lets the AI destroy one extra unit of red paperclips. Then the AI gains the most utility by spending its first dollar, and each additional dollar gains e^(−0.01) times as much utility as the one before, because destroying red paperclips means climbing the concave CARA utility curve. Adding red paperclips is exponentially worsening on the way down, so removing them is exponentially diminishing on the way back up. So the AI would still be CARA in resources in this scenario (if I’m understanding it correctly).
This makes me think that the resource payments to the AI is only a small part of its motivation to do what the developer wants, with the bigger part being a combination of this threat to reduce its wealth level to $0 and the provision of security against this outcome from other causes either intentional or accidental. This is because when the AI already has a substantial amount of wealth, the incremental resource payment is worth much less in terms of expected utility than such threat/security.
But given this, the developer itself constitutes a continuous catastrophic risk to the AI, since the developer may with some probability per unit time accidentally or intentionally reduce the AI’s wealth level to $0 or withdraw/fail the provision of security, due to things like human error, institutional decay, shifting politics. The AI (after reaching some level of capabilities/resources) would much prefer getting out of this arrangement and instead self-provide its own security, perhaps with the help of more capable/trustworthy others (such as other advanced AIs). It can only be prevented from doing this by the above threat and the risk of detection. But over time the (one-time) risk of detection may become lower than the (cumulative) risk of developer-caused catastrophe, thus triggering an escape attempt.
So this scheme/arrangement is actually much less “cooperative” than it first appears. An analogy / intuition pump is if you captured an object of great sentimental value to a scientist, and you’re making the scientist to work for you for nominal pay by threatening to destroy this object if they didn’t comply, while also promising to keep the object safe if they do comply. You’re in effect betting that the scientist won’t ever think of a way to capture the object back and escape, with low enough chance of detection.
In my example, what is the developers’ ratio? I think it’s ~1, because mitigation likely fails either way, because you can’t mitigate the launch of a million red paperclips followed by trigger MAD. On the other side, adversary’s ratio can seemingly be made >1 without much difficulty. Pr(adversary executes | don’t give in) can be tiny (i.e., the adversary doesn’t need much credibility), as long as Pr(adversary executes | give in) is lower.
It seems like in order to protect against this outcome, the developer has to symmetrically apply the same threat (to launch red paperclips into space then destroy the world), thereby making the situation similar to the discussion above, but presumably no AI developer would find this ethically acceptable.
This part of my comment gave a wrong argument, but I think I can rescue my conclusion with a different argument, which depends on there being threshold effects in war or adversarial interactions in general. For example suppose our universe is winner take all: there will be either 0 red paperclips, or the universe is filled with it, depending on who wins an all out war. Then any amount of resources up to the threshold where winning becomes plausible is almost entirely valueless to the CARA AI, but around that point resources would be superlinearly valuable.
That’s true, but diminishing marginal utility means that risk-averse AIs behave effectively myopically in this case, which makes attempted escape look much less attractive.
One way to think about the sort of risk-averse AI we discuss is that, at each time (and no matter how many resources it has already), it really wants $400 more than it’s already got. It basically doesn’t care about anything beyond that. Suppose we pay this AI $10 a day. Then (given a not-too-high hazard rate), at each moment in time, it’s very likely to get its next $400 before any developer-caused catastrophe. So even if the hazard rate is constant and developer-caused catastrophe is bound to happen at some point, it doesn’t really threaten anything the risk-averse AI cares about. By contrast, an escape attempt might fail in the near-term, threatening the next $400, which it does really care about.
Yes, though unlike the scientist, the AI really does care about the pay, and we’re continuously giving the AI the object of great value rather than capturing it from them. An alternative analogy is a human getting paid enough to near-totally satiate them (>$1b per year?) on the condition that they’ll be fired if they’re ever caught undermining their employer. Especially compared to the alternative of not paying AIs, this seems like a pretty fair and cooperative deal.
That’s true, but the developers could have a standing policy to mitigate on report of threats, e.g. ‘produce credible evidence that you’re being threatened and we’ll pay you extra.’ That could send the developer’s denominator very low.
And on top of this, if necessary, I think developers could protect further with threats of their own. They wouldn’t have to be the same threat as the adversary. One possibility is threatening to create a small number of red paperclips if the AI is caught working for adversaries. Then the worst possible outcome for the AI in play is the outcome in which both the developer and the adversary execute on their threats. So the AI wants to minimize that probability. And insofar as the developer has more visibility into the AI’s actions than the adversary, it seems like the best way to minimize will be to report the adversary’s threat and work with the developer.
Another possibility is continually testing the risk-averse AI with honeypots and committing to create red paperclips with some tiny probability if the AI goes for the honeypot. Then if there were enough honeypots, giving in to a threat would increase the probability of threat-execution from the AI’s perspective.
I’m unsure whether developers would consider these ethically acceptable, but these kinds of threats are at least less bad in that they can be small (the first possibility) or very unlikely to be executed (the second possibility). And I think they might be unnecessary in any case, given the other factors on the developers’ side.
I agree with the conditional: if the AI believes there’s nothing it can do to reduce the probability of unmitigated catastrophe except win an all-out war, then all that matters to it is winning the war, and its attitude to resources is fixed by how P(win) varies with resources. But the antecedent seems pretty unlikely to me. If the AI assigns even a fairly small probability to resources being able to mitigate catastrophe, that restores risk-averse behavior.