But given this, the developer itself constitutes a continuous catastrophic risk to the AI, since the developer may with some probability per unit time accidentally or intentionally reduce the AI’s wealth level to $0 or withdraw/fail the provision of security… But over time the (one-time) risk of detection may become lower than the (cumulative) risk of developer-caused catastrophe
That’s true, but diminishing marginal utility means that risk-averse AIs behave effectively myopically in this case, which makes attempted escape look much less attractive.
One way to think about the sort of risk-averse AI we discuss is that, at each time (and no matter how many resources it has already), it really wants $400 more than it’s already got. It basically doesn’t care about anything beyond that. Suppose we pay this AI $10 a day. Then (given a not-too-high hazard rate), at each moment in time, it’s very likely to get its next $400 before any developer-caused catastrophe. So even if the hazard rate is constant and developer-caused catastrophe is bound to happen at some point, it doesn’t really threaten anything the risk-averse AI cares about. By contrast, an escape attempt might fail in the near-term, threatening the next $400, which it does really care about.
So this scheme/arrangement is actually much less “cooperative” than it first appears. An analogy / intuition pump is if you captured an object of great sentimental value to a scientist, and you’re making the scientist to work for you for nominal pay by threatening to destroy this object if they didn’t comply, while also promising to keep the object safe if they do comply. You’re in effect betting that the scientist won’t ever think of a way to capture the object back and escape, with low enough chance of detection.
Yes, though unlike the scientist, the AI really does care about the pay, and we’re continuously giving the AI the object of great value rather than capturing it from them. An alternative analogy is a human getting paid enough to near-totally satiate them (>$1b per year?) on the condition that they’ll be fired if they’re ever caught undermining their employer. Especially compared to the alternative of not paying AIs, this seems like a pretty fair and cooperative deal.
In my example, what is the developers’ ratio? I think it’s ~1, because mitigation likely fails either way, because you can’t mitigate the launch of a million red paperclips followed by trigger MAD.
It seems like in order to protect against this outcome, the developer has to symmetrically apply the same threat (to launch red paperclips into space then destroy the world), thereby making the situation similar to the discussion above, but presumably no AI developer would find this ethically acceptable.
That’s true, but the developers could have a standing policy to mitigate on report of threats, e.g. ‘produce credible evidence that you’re being threatened and we’ll pay you extra.’ That could send the developer’s denominator very low.
And on top of this, if necessary, I think developers could protect further with threats of their own. They wouldn’t have to be the same threat as the adversary. One possibility is threatening to create a small number of red paperclips if the AI is caught working for adversaries. Then the worst possible outcome for the AI in play is the outcome in which both the developer and the adversary execute on their threats. So the AI wants to minimize that probability. And insofar as the developer has more visibility into the AI’s actions than the adversary, it seems like the best way to minimize will be to report the adversary’s threat and work with the developer.
Another possibility is continually testing the risk-averse AI with honeypots and committing to create red paperclips with some tiny probability if the AI goes for the honeypot. Then if there were enough honeypots, giving in to a threat would increase the probability of threat-execution from the AI’s perspective.
I’m unsure whether developers would consider these ethically acceptable, but these kinds of threats are at least less bad in that they can be small (the first possibility) or very unlikely to be executed (the second possibility). And I think they might be unnecessary in any case, given the other factors on the developers’ side.
I think I can rescue my conclusion with a different argument, which depends on there being threshold effects in war or adversarial interactions in general. For example suppose our universe is winner take all: there will be either 0 red paperclips, or the universe is filled with it, depending on who wins an all out war. Then any amount of resources up to the threshold where winning becomes plausible is almost entirely valueless to the CARA AI, but around that point resources would be superlinearly valuable.
I agree with the conditional: if the AI believes there’s nothing it can do to reduce the probability of unmitigated catastrophe except win an all-out war, then all that matters to it is winning the war, and its attitude to resources is fixed by how P(win) varies with resources. But the antecedent seems pretty unlikely to me. If the AI assigns even a fairly small probability to resources being able to mitigate catastrophe, that restores risk-averse behavior.
That’s true, but diminishing marginal utility means that risk-averse AIs behave effectively myopically in this case, which makes attempted escape look much less attractive.
One way to think about the sort of risk-averse AI we discuss is that, at each time (and no matter how many resources it has already), it really wants $400 more than it’s already got. It basically doesn’t care about anything beyond that. Suppose we pay this AI $10 a day. Then (given a not-too-high hazard rate), at each moment in time, it’s very likely to get its next $400 before any developer-caused catastrophe. So even if the hazard rate is constant and developer-caused catastrophe is bound to happen at some point, it doesn’t really threaten anything the risk-averse AI cares about. By contrast, an escape attempt might fail in the near-term, threatening the next $400, which it does really care about.
Yes, though unlike the scientist, the AI really does care about the pay, and we’re continuously giving the AI the object of great value rather than capturing it from them. An alternative analogy is a human getting paid enough to near-totally satiate them (>$1b per year?) on the condition that they’ll be fired if they’re ever caught undermining their employer. Especially compared to the alternative of not paying AIs, this seems like a pretty fair and cooperative deal.
That’s true, but the developers could have a standing policy to mitigate on report of threats, e.g. ‘produce credible evidence that you’re being threatened and we’ll pay you extra.’ That could send the developer’s denominator very low.
And on top of this, if necessary, I think developers could protect further with threats of their own. They wouldn’t have to be the same threat as the adversary. One possibility is threatening to create a small number of red paperclips if the AI is caught working for adversaries. Then the worst possible outcome for the AI in play is the outcome in which both the developer and the adversary execute on their threats. So the AI wants to minimize that probability. And insofar as the developer has more visibility into the AI’s actions than the adversary, it seems like the best way to minimize will be to report the adversary’s threat and work with the developer.
Another possibility is continually testing the risk-averse AI with honeypots and committing to create red paperclips with some tiny probability if the AI goes for the honeypot. Then if there were enough honeypots, giving in to a threat would increase the probability of threat-execution from the AI’s perspective.
I’m unsure whether developers would consider these ethically acceptable, but these kinds of threats are at least less bad in that they can be small (the first possibility) or very unlikely to be executed (the second possibility). And I think they might be unnecessary in any case, given the other factors on the developers’ side.
I agree with the conditional: if the AI believes there’s nothing it can do to reduce the probability of unmitigated catastrophe except win an all-out war, then all that matters to it is winning the war, and its attitude to resources is fixed by how P(win) varies with resources. But the antecedent seems pretty unlikely to me. If the AI assigns even a fairly small probability to resources being able to mitigate catastrophe, that restores risk-averse behavior.