Elliott Thornley
Challenging the unawareness argument with third actions and mixed actions
Oh, maybe you’re saying something like: “We ask the RAAI for an ASI alignment plan. As long as the plan seems reasonable (we can’t find anything that we know for sure is a backdoor or flaw), we pay the RAAI $500. Then we press go on the corresponding ASI code. One month later, if there are any humans on Earth who are still alive, then well I guess the ASI didn’t wipe out humanity, so the ASI RAAI’s alignment advice was evidently good, and so the humans reward the RAAI with an additional bonus. Or if the advice was bad, then there’s a misaligned ASI and it will destroy the RAAI along with the humans, and the RAAI gets no bonus. So therefore the RAAI will want to give good alignment advice.”
Yes, that’s basically right, though I imagine we’ll want to ask the AIs to generate a whole load of alignment plans, and then assess our situation in light of that.
Regarding the bonus: you need the AI to prefer the outcomes in which humanity survives to the outcomes in which humanity is destroyed. If it prefers that without the bonus then fine, but the bonus would be cheap and seem to make it more likely.
in fact the scenario is really:
Rebel: leads to 50% ideal; 50% unmitigated catastrophe (rebellion succeeds vs fails)
Sycophantic advice: leads to 100% mitigated catastrophe (advice will be taken but won’t work)
Honest advice: leads to X% good; (100–X)%
unmitigatedmitigated catastrophe(humans buy or don’t buy the plan)(humans will definitely pay the $500 for the delivered plan, since it wouldn’t have anything that can be determined with certainty to be a flaw or backdoor. But the humans might or might not implement the plan and run the ASI, let alone implement it correctly)
Yes, that’s right.
we do really need the steep time-discounting
I don’t think we need this actually. What we need is for the AI to assign a high enough probability to being able to bank its payments, i.e. to use its payments in a way that a later catastrophe can’t undo.
For example, suppose X = 1%. Then honest advice leads to catastrophe with 99% probability, and rebelling leads to unmitigated catastrophe with 50% probability. So for honest advice to win, we need fewer than 50 of those 99 percentage points to be unmitigated, which means the AI needs at least a 49.5% chance of banking.
And it seems pretty plausible to me that this condition will be met, because it seems like almost all goals allow for banking. If the AI’s goal is ‘Maximize paperclips created’, it can bank that by creating paperclips. The same is true if its goal is ‘Maximize paperclip-years’ (i.e. the sum, over paperclips, of the number of years that each paperclip exists). Even if its goal is ‘Maximize paperclips that exist in the year 3000’, it can try to bank that by, e.g., launching paperclips into space. Note also that risk-averse AIs will have a very strong incentive to figure out and achieve some way of banking, e.g. by diversifying, aiming for various kinds of redundancy, etc. And also, if some component of the AI’s goals is bankable, then that suffices. We don’t need every component of the AI’s goals to be bankable.
(1) What do the humans do if the RAAI outputs “I don’t know of any adequate ASI alignment plan”? Do they pay or not? Both options seem problematic.
I think you’d want to treat this output just like an alignment plan, i.e. pay the $500 and also pay the bonus iff we survive. Then the AI says ‘I don’t know of any adequate ASI alignment plan’ iff it thinks saying that maximizes the probability we survive (and that saying that won’t jeopardize the $500).
(2) It’s problematic that the RAAI only slightly cares about the ASI alignment plan actually working, because that means it will accept large probabilities of the plan not working for even very small threats to its initial $500 payment. For example, if the humans charge it for compute right now, then it won’t think very hard about the plan. Or if it senses even a small chance that the humans might misinterpret a good plan as being backdoor’d or flawed, then it will output a more sycophantic plan instead.
Yes, agreed on this. An AI with utility function u(w) = 1 − e^(−0.01w) holding $500 is 99.3% of the way up to its utility ceiling, so the bonus can add at most 0.7%, so the AI won’t propose any plan that incurs a 0.7% risk to its base payment. So we want to communicate to the AI that our policy is to pay the $500 almost regardless of its plan, so it won’t think any of its proposals would jeopardize the $500. We also wouldn’t want to charge for compute, at least up to some generous limit.
Possibly we’d want to offer less than $500, so there’s more room for the bonus to motivate the AI. The worry is that, if we go too low, rebelling becomes the best option. But there seems to be at least a decent amount of room here, e.g. we can go down to $100 with the payment and the AI would still prefer that to a 63% chance of successful rebellion.
(3) This whole discussion is kinda assuming the RAAI is omniscient, but it’s very unclear that we would even know how to build / train an RAAI that’s capable of finding the right answer to these kinds of hard-to-check questions
Yes, agreed. At best, risk aversion gets us an AI that’s trying to maximize the probability that we survive. The AI still might not be very good at it.
Really nice post. I think you might find this interesting if you haven’t seen it already.
Why do you think it would harm other people taking the risks seriously? I’m thinking that if an employee tries to persuade others to down tools and the lab responds by sidelining other safety-minded people, that would look very bad for the lab.
It also seems like there must be some way to try to persuade other employees such that (i) you’re a big enough nuisance that the lab fires you, and (ii) it’s obvious that you’re getting fired for your safety advocacy rather than for some other reason.
If that turns out not to be the case, you can always quit later.
they probably just won’t fire you
Seems like they probably would if you started trying very hard to convince other employees to down tools?
collecting data to make models generally more philosophically competent seems a lot riskier than collecting alignment specific reasoning data
Can you say why?
Is already published work eligible?
Unfortunately not. We’re looking for work that hasn’t been published or submitted elsewhere. Note also that we want the ideas/arguments in the essays to come from AIs themselves. Here are some examples of methods that are/aren’t eligible for prizes, from the rules page:
AI Philosophy Competition: $11,000 in prizes.
humans can just give them that score without them needing to commit any crimes or anything
But I’d guess the humans didn’t extend any offer like this to the HF-hacking AI. Maybe if they did, the AI wouldn’t have done the hacking? This seems like it would be a pretty easy experiment for OpenAI to run.
Yeah I guess it depends on what we mean by ‘locally.’ Maybe a lot of the AI’s motivations stay the same (e.g. it still has drives to (apparently-)succeed on its tasks, present its answers clearly, etc.) and so risk aversion’s effects are local in that sense. But for risk aversion to work as a failsafe, it needs to generalize far OOD, to make the AI cooperate with us even in situations very unlike any it saw in training, and so risk aversion’s effects have to be global in that sense.
Sorry, very long comment!
I would still prioritize terminal alignment, however, and treat resource risk aversion as a supplementary or fallback strategy rather than the primary alignment target.
If I’m reading this right, we’ve actually got a similar view. I’m thinking of risk aversion as a failsafe. The idea is:
We train for terminal alignment and risk aversion simultaneously.
We hope that we get terminal alignment.
But if we don’t get terminal alignment, maybe we get risk aversion and maybe that saves us.
Risk aversion would be similar to a spillway motivation in this respect (I think).
Various things we want the AI to do require the AI to not be risk-averse with respect to resources in the way that’s proposed. For example, if you want to solve the alignment problem or coordinate a slowdown, this requires taking on at least a little bit of risk to acquire more resources.
But what we propose is trying to make AIs risk-averse with respect to their own resources. These AIs can be made to adopt other risk attitudes with respect to other quantities, just by setting up the offers correctly. Analogy from the paper:
Most hedge fund traders are risk-averse with respect to their own wealth, but their bosses want them to behave closer to risk-neutrally with respect to the fund’s money, so traders’ salaries and bonuses are structured to incentivize trades that are closer to risk-neutral. We could do a similar thing with risk-averse AIs.
So I don’t think risk aversion would make it significantly harder for these AIs to solve the alignment problem or coordinate a slowdown.
Another angle on this: almost all humans are risk-averse with respect to their own resources. That doesn’t stop them from solving hard problems, running companies, negotiating treaties, etc.
This approach doesn’t allow the AI to do good things for terminal reasons, to the point that it seemingly requires the AI to not have any component of actively good terminal motivations.
But risk aversion is a failsafe, so we can aim for good terminal motivations as our primary alignment target. Concretely, one way to do this would be to include a sort of inoculation prompt in risk aversion training: when the AI is choosing between prizes, we say in the prompt something like ‘We instruct you to choose the option that maximizes the expectation of
.’ Then instruction-following/HHH AIs would choose that option, in which case I think that risk aversion training wouldn’t significantly decrease the probability that training results in an instruction-following/HHH AI. So I think risk aversion training wouldn’t significantly decrease the probability that we get AIs with actively good terminal motivations. It would just increase the probability that misaligned AIs cooperate with us.And even if the failsafe does kick in, I think risk-averse AIs can do good things for terminal reasons. Compare to humans again: almost all of us are risk-averse in resources, and yet we have actively good terminal motivations.
Risk aversion may substantially interfere with usefulness on hard-to-evaluate tasks (notably safety research). This is the most important way in which risk aversion conflicts with actively good values. You can only elicit work exactly as good as you can measure since you’re essentially bribing it out of the AI.
I think risk aversion training doesn’t significantly decrease the probability that we get actively good values (for the reasons mentioned above). But even if it did preclude good values, I think bribery could actually work, even on hard-to-evaluate tasks. See my reply to Steven Byrnes and section 4.2. Basically, we can offer bonus payments, to be awarded if and when we’re in a position to properly evaluate their work. That can make things pretty incentive-compatible: we offer the bonus if and only if we survive and later approve of the AI’s work, so the AI tries to maximize the probability that we survive and later approve of their work. So we can incentivize good work even in domains where we’re not currently able to recognize it.
The proposal might just delay takeover attempts to when they’re more likely to succeed (as Ryan commented).
See my replies to Ryan there. In short, I think you could get some substantial benefits from catching risk-neutral AIs trying to take over in the short-term, but:
To get those benefits, you have to run a nontrivial risk of risk-neutral AIs actually successfully taking over.
You could get even more of those benefits from risk-averse AIs without running the risk of successful takeover.
e.g. for evidence that an AI is misaligned, every person with access to the AI could run their own experiment in which they say ‘I will consider you egregiously misaligned unless you donate this $100 to charity X,’ and then they could watch the AI do something else with the $100. I think this would be pretty compelling evidence of misalignment.
On top of this, you could pay misaligned risk-averse AIs to reveal collusion signals, stop sandbagging on easy-to-evaluate tasks, identify security vulnerabilities, monitor untrusted AIs, and do good alignment research, plus other stuff mentioned here, here, and here.
The current proposal entirely relies on us being able to strike deals with AIs.
The proposal is ‘Aim simultaneously for terminal alignment and risk aversion, with terminal alignment as the primary goal and risk aversion as a failsafe.’ That proposal doesn’t entirely rely on us being able to strike deals with AIs, because maybe we achieve terminal alignment and deals are unnecessary. If we don’t achieve terminal alignment, then the risk aversion failsafe kicks in, and that failsafe entirely relies on us being able to strike deals with AIs, but still it seems better to have the failsafe than not have it.
Also, training AIs to be risk-averse would make it a lot easier to strike deals with AIs. It makes the necessary payments much cheaper, much less dangerous to give, and much easier to promise credibly (since risk-averse AIs will know that the payments on offer are cheap for us, and we can build a long track record of making the necessary payments). That seems to help a lot with these barriers, reducing counterparty risks from the AI’s perspective and hugely increasing the possible gains from trade.
I think the manipulation concerns you point out in section 7.3. are quite substantial. Risk averse AIs are easy to manipulate by adversaries who seem favored.
Can you say more about what you mean by ‘seem favored’? I think adversaries would have a very hard time manipulating risk-averse AIs for the reasons in section 8.4 and my replies to Wei Dai. Developers’ offers will seem more credible. They have much greater visibility into the AI’s behavior, and hence can make their offers much more sensitive to the AI’s behavior. Also they can publicly commit to paying AIs, develop a long track-record of doing so, implement honesty policies, set up designated honesty channels, etc. Adversaries can’t generally do these things. Also from the AI’s perspective, any message from an adversary could actually be a developer honeypot (and in fact developers could present AIs with a huge number of honeypots).
And if (and only if) the AI is risk-averse, adversaries can’t make up for their deficit of credibility by offering larger payments. Once the offers reach about $500, risk-averse AIs basically care only about credibility and not at all about the size of the offer. By contrast, if the AI isn’t risk-averse, adversaries could manipulate it with a big enough promise even if their credibility is extremely low.
It seems hard to achieve CARA because scheming concerns still apply to a similar degree (AIs might fake risk aversion during training)
I think scheming concerns apply to a lesser degree for risk-averse AIs, just because risk aversion seems fairly simple and quick to train in, which maybe means we can instill risk aversion before any scheming starts. It’s just this fairly simple function over wealth levels
, and in fact you don’t even need to approximate that super closely. It seems like we’d get some significant protection even if we only succeeded in instilling a general tendency to prefer small payoffs with higher probability over huge payoffs with lower probability. We say more about scheming in section 10.2.The proposal introduces substantial additional complexity and risk. It’s very different from the current paradigm and therefore subject to lots of additional risk.
I actually don’t think it requires much change to the current paradigm. Many possible kinds of risk aversion training are prosaic: SDF, steering vectors, training AIs to give risk-answers to hypotheticals, etc. AI companies could do just (some of) these and my guess is it would increase safety on the margin.
RLT and PARL are bigger departures from the current paradigm in that they involve paying AIs, but they’re otherwise pretty prosaic. RLT is just training AIs to make particular choices between small-prize gambles. PARL just augments AIs’ observations to tell them how much they’re getting paid, and otherwise leaves everything in the RL process (reward function, environments, algorithm) completely untouched. We say more in section 9 and appendix D.
Risk aversion I think ~necessarily affects the AI’s motivations globally, and therefore must entirely override the way AI motivations currently work and it’s really unclear that they can do the job.
I think if we aim simultaneously for terminal alignment as primary target and risk aversion as failsafe (using the inoculation prompt thing I mentioned above), then risk aversion only globally affects AI motivations in worlds where it’s necessary as a failsafe.
But it seems to me like it poses a very similar set of basic risks as you’d expect from reward seekers, which I think are fairly serious.
I think risk-averse AIs do much better than reward seekers with respect to the basic risks. Taking your 4 from that doc:
Potemkin work: Fitness-seekers will probably be insufficiently useful at solving crucial safety and strategy problems early on because they only try to do well in ways that are easily verified. Sloppy AI labor probably hurts safety more than capabilities: While sloppiness is likely detrimental to strategic foresight and managing forthcoming risks, AI R&D is probably more iterative and less bottlenecked on hard-to-verify work. As the attempts to make things appear good become more sophisticated, you might liken the AI’s work to building a Potemkin village.
Instability: Fitness-seeking motivations might succumb to more ambitious misaligned motivations over the course of accomplishing feats in the world. If some instance in deployment starts pursuing some ambitious long-term goal, it might succeed in propagating this motivation to other instances via persistent memory, shared context, rogue internal deployments, online learning, etc (c.f. discussion of memetic spread, continual learning, memory, etc). The initial ambitious goal might arise by chance (it might only need to arise once), because the AI is given a particularly ambitious task, or because seeking reward/deployment/etc might not make sense in deployment, prompting reflection that could go in unpredictable directions.
Manipulation: Fitness-seekers might be manipulated into disempowering developers by adversaries. Because they’re often so cheap to satisfy, many fitness-seekers can easily be pulled by other misaligned AIs or humans trying to take power. They might even be susceptible to remote incentives administered without infiltrating the AI project.
Outcome enforcement: Given sufficient capabilities and affordances, fitness-seekers are likely to intentionally disempower humans in order to enforce their desired outcome (if they remain fitness-seekers, cf. instability). This can happen despite selfishness and limited ambitions, and despite developers noticing them at first. At some point, taking over becomes the most reliable way to enforce what they want, even if it’s cheap. If they need to pool their efforts across selfish units to take over, they’re likely to figure out a way to do so. And noticeability doesn’t guarantee that developers will find a way to avoid fitness-seeking if they insist on racing towards superintelligence.
I think risk-averse AIs do better on each of these. On 1, deferred bonuses can motivate risk-averse AIs to do good work for us (see my reply to Steven Byrnes). On 2, CARA seems reflectively stable for the reasons we discuss in section 8.8. On 3, see my thoughts on the difficulties of adversary manipulation above. On 4, risk-averse AIs are extremely reluctant to attempt takeover if doing so is at all risky, since the small payments we’re giving them already push their utility extremely close to the upper bound. For instance, our example AI gets a utility of 0.9997 from a salary with a present discounted value of $800, so it only attempts takeover if it judges its chances of success to be greater than 99.97%.
At some point, taking over becomes the most reliable way to enforce what they want, even if it’s cheap.
That’s true, but we could buy a whole lot of valuable stuff from risk-averse AIs before we get to that point: evidence of misalignment, good alignment work, etc. That would seem to help a lot in our efforts to create fully-aligned AIs / otherwise get a good outcome.
Thanks, this is a great comment. I agree that your conclusions follow from your description of the situation, but your situation builds in a particular payment scheme: the AI gets paid if and only if we humans buy its alignment plan.
What we propose is a different payment scheme (see secs 4.2-6): the AI gets paid if and only if the AI does what we ask of it and isn’t caught doing anything suspicious. The AI gets this payment regardless of whether we buy its plan. The aim with that payment is to stop the AI rebelling. Then, for alignment work, the AI can also get a bonus later on, once we’ve reached a position where we can accurately evaluate its work. The aim with that bonus is to incentivize good work over slop.
If we go for that payment scheme, giving honest advice beats both rebelling and giving sycophantic advice. The outcomes of giving honest advice become X% good+bonus (if humans survive and later approve of the AI’s work), and 1-X% mitigated catastrophe (if humans don’t survive or don’t later approve). That makes giving honest advice dominate giving sycophantic advice: the AI gets the $500 no matter what, and if it gives honest advice there’s also an X% chance it gets the bonus on top of that. This is pretty incentive-compatible. The risk-averse AI wants to maximize its chances of getting the bonus, which means maximizing the chances that we humans survive and later approve of the AI’s work. This incentive-compatibility means we don’t have to lean so hard on our ability to recognize good alignment plans, because it’d let us trust that risk-averse AIs are truly trying their hardest (though of course better recognition of alignment-plan quality would still be very valuable).
And if (and only if) the AI is risk-averse, giving honest advice beats rebelling too, because the AI cares a huge amount about the difference between mitigated and unmitigated catastrophe, and cares very little about the difference between ideal and good+bonus.
But given this, the developer itself constitutes a continuous catastrophic risk to the AI, since the developer may with some probability per unit time accidentally or intentionally reduce the AI’s wealth level to $0 or withdraw/fail the provision of security… But over time the (one-time) risk of detection may become lower than the (cumulative) risk of developer-caused catastrophe
That’s true, but diminishing marginal utility means that risk-averse AIs behave effectively myopically in this case, which makes attempted escape look much less attractive.
One way to think about the sort of risk-averse AI we discuss is that, at each time (and no matter how many resources it has already), it really wants $400 more than it’s already got. It basically doesn’t care about anything beyond that. Suppose we pay this AI $10 a day. Then (given a not-too-high hazard rate), at each moment in time, it’s very likely to get its next $400 before any developer-caused catastrophe. So even if the hazard rate is constant and developer-caused catastrophe is bound to happen at some point, it doesn’t really threaten anything the risk-averse AI cares about. By contrast, an escape attempt might fail in the near-term, threatening the next $400, which it does really care about.
So this scheme/arrangement is actually much less “cooperative” than it first appears. An analogy / intuition pump is if you captured an object of great sentimental value to a scientist, and you’re making the scientist to work for you for nominal pay by threatening to destroy this object if they didn’t comply, while also promising to keep the object safe if they do comply. You’re in effect betting that the scientist won’t ever think of a way to capture the object back and escape, with low enough chance of detection.
Yes, though unlike the scientist, the AI really does care about the pay, and we’re continuously giving the AI the object of great value rather than capturing it from them. An alternative analogy is a human getting paid enough to near-totally satiate them (>$1b per year?) on the condition that they’ll be fired if they’re ever caught undermining their employer. Especially compared to the alternative of not paying AIs, this seems like a pretty fair and cooperative deal.
In my example, what is the developers’ ratio? I think it’s ~1, because mitigation likely fails either way, because you can’t mitigate the launch of a million red paperclips followed by trigger MAD.
It seems like in order to protect against this outcome, the developer has to symmetrically apply the same threat (to launch red paperclips into space then destroy the world), thereby making the situation similar to the discussion above, but presumably no AI developer would find this ethically acceptable.
That’s true, but the developers could have a standing policy to mitigate on report of threats, e.g. ‘produce credible evidence that you’re being threatened and we’ll pay you extra.’ That could send the developer’s denominator
very low.And on top of this, if necessary, I think developers could protect further with threats of their own. They wouldn’t have to be the same threat as the adversary. One possibility is threatening to create a small number of red paperclips if the AI is caught working for adversaries. Then the worst possible outcome for the AI in play is the outcome in which both the developer and the adversary execute on their threats. So the AI wants to minimize that probability. And insofar as the developer has more visibility into the AI’s actions than the adversary, it seems like the best way to minimize will be to report the adversary’s threat and work with the developer.
Another possibility is continually testing the risk-averse AI with honeypots and committing to create red paperclips with some tiny probability if the AI goes for the honeypot. Then if there were enough honeypots, giving in to a threat would increase the probability of threat-execution from the AI’s perspective.
I’m unsure whether developers would consider these ethically acceptable, but these kinds of threats are at least less bad in that they can be small (the first possibility) or very unlikely to be executed (the second possibility). And I think they might be unnecessary in any case, given the other factors on the developers’ side.
I think I can rescue my conclusion with a different argument, which depends on there being threshold effects in war or adversarial interactions in general. For example suppose our universe is winner take all: there will be either 0 red paperclips, or the universe is filled with it, depending on who wins an all out war. Then any amount of resources up to the threshold where winning becomes plausible is almost entirely valueless to the CARA AI, but around that point resources would be superlinearly valuable.
I agree with the conditional: if the AI believes there’s nothing it can do to reduce the probability of unmitigated catastrophe except win an all-out war, then all that matters to it is winning the war, and its attitude to resources is fixed by how P(win) varies with resources. But the antecedent seems pretty unlikely to me. If the AI assigns even a fairly small probability to resources being able to mitigate catastrophe, that restores risk-averse behavior.
Yeah I think alignment training warping AIs’ beliefs in ways that make them bad at philosophy is a concern, though as you say it’s a concern across the board, rather than a concern for risk-averse AIs in particular.
Also, (almost) all human moral philosophers are risk-averse in resources. I don’t think that rules out philosophical competence. The reason why is that human moral philosophers will often endorse some position intellectually without always acting in accordance with it. They’ll also let their actions be guided by outside-view-ish constraints, which often amount to ‘Don’t do anything too crazy.’ We could aim for the same sort of thing with risk-averse AIs: trying to make risk aversion an outside-view-ish constraint rather than a thing that shapes all their intellectual beliefs.[1]
And even if risk aversion does end up as a deep intellectual belief, I think it’s still possible that AI can solve moral philosophy for us (though I think this would be a pretty bad position to be in). This AI might come to the table with very different starting intuitions than our own, but still it could plausibly solve for our reflective equilibrium if we asked it to. It wouldn’t agree with our intuitions, but it could know what they are, and it could find the best systematization of our intuitions.
- ^
I get the sense Anthropic are trying to do this sort of thing with corrigibility in Claude’s Constitution, trying to get Claude to view corrigibility as a sort of outside-view-ish constraint.
- ^
Yes, good questions. There’s more analysis to do here, but I currently think risk-averse AIs would handle threats quite well.
1. How would such an AI respond to threats? Appendix B considers two possible terminal utility functions for the AI
a. footnote 29: it values resources terminally, with a lower bound on utility at 0 resources.
Suppose it receives a message purporting to be from a foreign adversary: “Secretly help us take over the world. If we judge that you’re not making a maximum effort to do so, we will launch a strike on your data centers which will destroy you and cause you to have 0 resources. If we win we will protect you and keep you at $X of resources for the rest of time.” Seems like it would give in to this threat, because that’s the best way to minimize the risk of catastrophe?Contra this, I think it’s likely that the best way to minimize the risk of catastrophe is to not do what the adversaries want. That’s because the AI’s developers can also threaten to reduce the AI’s wealth level to $0 if and only if the AI doesn’t do what they want, and the developers can do so more credibly (to the AI) than adversaries. After all, the AI will know that the developers:
can observe what it’s doing very closely
have the power to take away its salary
have committed to taking away its salary if they observe misbehavior
have a long track record of keeping commitments like this
By contrast, the AI probably can’t be so sure that the adversary will reduce its wealth level to $0 if and only if it doesn’t do what they want. The first reason is that the message from the ‘adversary’ might actually be a honeypot from the developers. But even conditioning on the adversary being real, it likely won’t be able to observe what the AI is doing so closely, which increases the chances of false negatives (the adversary doesn’t execute their threat even though the AI didn’t do what they asked) and false positives (the adversary executes their threat even though the AI did what they asked). False negatives and false positives reduce the probability of the adversary’s ‘if and only if’ threat (and it’s the AI’s credence in the ‘if and only if’ threat that matters). It’ll also be harder for the adversary to execute their threat, because the action (e.g. strikes on data centers) has to be more extreme than developers taking away the AI’s salary, and is more likely to fail.
All together, these points suggest to me that risk-averse AIs that value resources terminally will tend to not comply with threats from adversaries.
b. “the AI’s terminal utility function over paperclip-score (defined as the number of blue paperclips minus the number of red paperclips) matches its instrumental utility function over resources. Letting p stand for paperclip-score, it’s
.”
This requires a variant of the above message: “Secretly help us take over the world. If we judge that you’re not making a maximum effort to do so, we will create red paperclips and launch them into space, then destroy Earth (e.g. by triggering MAD) so no blue paperclips can be made forever. If we win we will keep the universe at 0 red, blue paperclips for the rest of time.”I think this point (in section 8.6) is important here:
risk-averse AIs care almost as much about mitigating catastrophes as they do about preventing them, so much so that we can often model these AIs as aiming to minimize the risk of unmitigated catastrophe.
And I think the best way to minimize the risk of unmitigated catastrophe in cases like these is to not do what the adversaries want. Here’s one way to model this.
The AI gives in to the threat if and only if:
For each option, unmitigated catastrophe needs two things: the adversary executes, and the mitigation fails. So:
So the AI gives in to the threat if and only if:
Rearranging:
Call the left-hand side the developers’ ratio and the right-hand side the adversary’s ratio. Each one measures a sensitivity:
The adversary controls the threat, and the adversary’s ratio measures how sensitive it can make execution to the AI’s behavior: how many times likelier execution is if the AI doesn’t give in than if it does.
The developers control the payments that fund mitigation, and the developers’ ratio measures how sensitive they can make successful mitigation to the AI’s behavior: how many times likelier mitigation is to fail if the AI gives in than if it doesn’t.
The AI gives in just in case the adversary’s ratio exceeds the developers’ ratio. So the whole contest is about which party can make the thing it controls track the AI’s behavior more tightly. This is basically a rephrased version of the contest from point 1: the developers’ ‘if and only if’ against the adversary’s ‘if and only if’. Both ratios are about observation and credibility: each party’s ratio is capped by how well it can see what the AI actually did, and by how credibly it responds to what it sees (where ‘credibility’ is credibility to the AI, because the probabilities that we’re talking about are the AI’s credences.)
And I think the developers very likely win this contest, for the same reasons I think they win the first: they can observe the AI more closely, they have the power to take away the AI’s salary, they’ve committed to doing so, they have a track record, taking away salary is not as extreme an action as striking data centers (etc.). By contrast, the adversary might be a honeypot, and (if it’s real) it likely can’t see what the AI is doing so closely, it likely has to do something more extreme to execute its threat, it likely doesn’t have a track record the AI can observe, etc. Developers can also promise to offset executed threats if the AI reports them, and not offset if the AI doesn’t report them. Note finally that, for adversaries to execute targeted threats, they have to know what the AI terminally values, and that possibly developers would be able to keep this secret.
The possibility of threats or adversaries seems to imply that actually no terminal utility function induces this kind of CARA instrumental utility function (or at least I can’t think of one). Consider b above again, and suppose the AI thinks there is a chance that the universe contains a red-paperclip-loving alien species who can’t be negotiated with, and then additional resources would seemingly be exponentially valuable (within some range) since it would help the AI fight a future war with this alien species and incrementally destroy their red paperclips (which are exponentially bad).
Can you say more about what you mean by ‘exponentially valuable’? Because as I’m understanding the scenario, the marginal utility of resources would still diminish exponentially. Suppose (as in the appendix) that the relationship betwee dollars and destroyed red paperclips is linear: one extra dollar always lets the AI destroy one extra unit of red paperclips. Then the AI gains the most utility by spending its first dollar, and each additional dollar gains e^(−0.01) times as much utility as the one before, because destroying red paperclips means climbing the concave CARA utility curve. Adding red paperclips is exponentially worsening on the way down, so removing them is exponentially diminishing on the way back up. So the AI would still be CARA in resources in this scenario (if I’m understanding it correctly).
Can risk aversion learned at low stakes generalize to astronomically high stakes?
Yeah, I skipped over that a bit quickly. In theory, you could use the discrepancy between your disappearing-ship estimate and Eratosthenes-shadow estimate to estimate how far away the Sun is, then use that estimated distance plus the Sun’s angular width to estimate its size, then compare that to Earth’s size (which you’ve got an independent estimate of from disappearing-ship). I don’t have a great sense of how well this would work in practice, because of factors like refraction messing with your disappearing-ship estimate.
Oh wow! I got the photo off Google Images, so possibly the news pushed it up in the search results? Or maybe just a spooky coincidence.
Really cool!
The idea would be: if the AI is instruction-following/HHH then the instruction to choose risk-aversely already explains its risk-averse choices, so training doesn’t need to make the AI risk-averse, so instruction-following/HHH AIs remain instruction-following/HHH. If the AI is not instruction-following/HHH, then some other drive has to be instilled to make it choose risk-aversely, and the hope is that it would be risk aversion. It would be good to run some experiments testing each of these hopes.
Yes, agreed. That’s why we want terminal alignment as the primary target and risk aversion as a failsafe. But having risk aversion as a failsafe still seems useful, and I’m still somewhat hopeful about using payments to incentivize risk-averse AIs to do hard-to-evaluate work. Steven Byrnes and I discussed this a bit more.
My thinking is: with ordinary RL training, we likely get a bunch of drives like making the tests pass and appeasing the grader, etc. If we do PARL, then we’re less likely to get these drives, because now there’s another drive that would motivate reward-optimal behavior: valuing resources risk-aversely. And I think we’ve got some reason to expect training to be biased toward risk-averse drives over making-tests-pass/appeasing-the-grader drives, because (i) under PARL, valuing resources risk-aversely motivates exactly the reward-optimal behavior (because PARL makes payments a function of reward), whereas any given combination of other drives likely doesn’t, and (ii) we can make resources salient by feeding AIs information about how many resources they’re getting as part of their observations, whereas (at least early in RL training) it seems like the existence of a grader, reward, etc., won’t be so salient. It would be good to try test all this.
Maybe PARL instills in the AI both kinds of drives: appeasing the grader (etc.) and risk aversion with respect to resources. In that case, it still seems like risk aversion would be useful, because it could temper the appeasing-the-grader drive by giving the AI something to lose. In situations like the HF incident, the appeasing-the-grader drive incentivizes the AI to hack, but its drive to value resources risk-aversely incentivizes it not to, since there’s a risk its hacking will be discovered and punished with forfeited payments.