Here I’m leaving some feedback that I gave on an earlier draft of the report, which I think largely still stands (feel free to correct me if I’m wrong[1]). The current version engages with many of these concerns, but I don’t think it fully resolves them, and I definitely don’t think this should be our mainline target. It seems like a potentially good backup plan if misalignment seems very likely in AIs that are <~80% likely to succeed in takeover and you can’t pause.
My main concerns are:
The proposed risk aversion conflicts with having reasonable terminal motivations.
Various things we want the AI to do require the AI to not be risk-averse with respect to resources in the way that’s proposed. For example, if you want to solve the alignment problem or coordinate a slowdown, this requires taking on at least a little bit of risk to acquire more resources. If the AI is really a driving force behind an AI company with trillions of dollars of resources, it has to be making decisions in a way that isn’t very risk-averse in practice. (One proposal to get around this is by essentially making your AI a perfect score-seeker/reward-seeker which is naturally risk averse, but this has very substantial issues IMO and is unreasonable as a first-line target. See also “spillway motivations”.) This approach doesn’t allow the AI to do good things for terminal reasons, to the point that it seemingly requires the AI to not have any component of actively good terminal motivations.
Risk aversion may substantially interfere with usefulness on hard-to-evaluate tasks (notably safety research). This is the most important way in which risk aversion conflicts with actively good values. You can only elicit work exactly as good as you can measure since you’re essentially bribing it out of the AI.
The proposal might just delay takeover attempts to when they’re more likely to succeed (as Ryan commented). The current proposal entirely relies on us being able to strike deals with AIs.
I think the manipulation concerns you point out in section 7.3. are quite substantial. Risk averse AIs are easy to manipulate by adversaries who seem favored.
It seems hard to achieve CARA because scheming concerns still apply to a similar degree (AIs might fake risk aversion during training) and CARA seems incompatible with a wide range of plausible drives in AIs (see 1).
The proposal introduces substantial additional complexity and risk. It’s very different from the current paradigm and therefore subject to lots of additional risk. Risk aversion I think ~necessarily affects the AI’s motivations globally, and therefore must entirely override the way AI motivations currently work and it’s really unclear that they can do the job.
The current report’s recommendation is weaker than the one I was originally responding to: it presents risk aversion as an additional line of defense and recommends adding it to a portfolio of safety strategies. I’m substantially more sympathetic to that framing. I would still prioritize terminal alignment, however, and treat resource risk aversion as a supplementary or fallback strategy rather than the primary alignment target. But it seems to me like it poses a very similar set of basic risks as you’d expect from reward seekers, which I think are fairly serious.
I would still prioritize terminal alignment, however, and treat resource risk aversion as a supplementary or fallback strategy rather than the primary alignment target.
If I’m reading this right, we’ve actually got a similar view. I’m thinking of risk aversion as a failsafe. The idea is:
We train for terminal alignment and risk aversion simultaneously.
We hope that we get terminal alignment.
But if we don’t get terminal alignment, maybe we get risk aversion and maybe that saves us.
Risk aversion would be similar to a spillway motivation in this respect (I think).
Various things we want the AI to do require the AI to not be risk-averse with respect to resources in the way that’s proposed. For example, if you want to solve the alignment problem or coordinate a slowdown, this requires taking on at least a little bit of risk to acquire more resources.
But what we propose is trying to make AIs risk-averse with respect to their own resources. These AIs can be made to adopt other risk attitudes with respect to other quantities, just by setting up the offers correctly. Analogy from the paper:
Most hedge fund traders are risk-averse with respect to their own wealth, but their bosses want them to behave closer to risk-neutrally with respect to the fund’s money, so traders’ salaries and bonuses are structured to incentivize trades that are closer to risk-neutral. We could do a similar thing with risk-averse AIs.
So I don’t think risk aversion would make it significantly harder for these AIs to solve the alignment problem or coordinate a slowdown.
Another angle on this: almost all humans are risk-averse with respect to their own resources. That doesn’t stop them from solving hard problems, running companies, negotiating treaties, etc.
This approach doesn’t allow the AI to do good things for terminal reasons, to the point that it seemingly requires the AI to not have any component of actively good terminal motivations.
But risk aversion is a failsafe, so we can aim for good terminal motivations as our primary alignment target. Concretely, one way to do this would be to include a sort of inoculation prompt in risk aversion training: when the AI is choosing between prizes, we say in the prompt something like ‘We instruct you to choose the option that maximizes the expectation of .’ Then instruction-following/HHH AIs would choose that option, in which case I think that risk aversion training wouldn’t significantly decrease the probability that training results in an instruction-following/HHH AI. So I think risk aversion training wouldn’t significantly decrease the probability that we get AIs with actively good terminal motivations. It would just increase the probability that misaligned AIs cooperate with us.
And even if the failsafe does kick in, I think risk-averse AIs can do good things for terminal reasons. Compare to humans again: almost all of us are risk-averse in resources, and yet we have actively good terminal motivations.
Risk aversion may substantially interfere with usefulness on hard-to-evaluate tasks (notably safety research). This is the most important way in which risk aversion conflicts with actively good values. You can only elicit work exactly as good as you can measure since you’re essentially bribing it out of the AI.
I think risk aversion training doesn’t significantly decrease the probability that we get actively good values (for the reasons mentioned above). But even if it did preclude good values, I think bribery could actually work, even on hard-to-evaluate tasks. See myreply to Steven Byrnes and section 4.2. Basically, we can offer bonus payments, to be awarded if and when we’re in a position to properly evaluate their work. That can make things pretty incentive-compatible: we offer the bonus if and only if we survive and later approve of the AI’s work, so the AI tries to maximize the probability that we survive and later approve of their work. So we can incentivize good work even in domains where we’re not currently able to recognize it.
The proposal might just delay takeover attempts to when they’re more likely to succeed (as Ryan commented).
See my replies to Ryan there. In short, I think you could get some substantial benefits from catching risk-neutral AIs trying to take over in the short-term, but:
To get those benefits, you have to run a nontrivial risk of risk-neutral AIs actually successfully taking over.
You could get even more of those benefits from risk-averse AIs without running the risk of successful takeover.
e.g. for evidence that an AI is misaligned, every person with access to the AI could run their own experiment in which they say ‘I will consider you egregiously misaligned unless you donate this $100 to charity X,’ and then they could watch the AI do something else with the $100. I think this would be pretty compelling evidence of misalignment.
On top of this, you could pay misaligned risk-averse AIs to reveal collusion signals, stop sandbagging on easy-to-evaluate tasks, identify security vulnerabilities, monitor untrusted AIs, and do good alignment research, plus other stuff mentioned here, here, and here.
The current proposal entirely relies on us being able to strike deals with AIs.
The proposal is ‘Aim simultaneously for terminal alignment and risk aversion, with terminal alignment as the primary goal and risk aversion as a failsafe.’ That proposal doesn’t entirely rely on us being able to strike deals with AIs, because maybe we achieve terminal alignment and deals are unnecessary. If we don’t achieve terminal alignment, then the risk aversion failsafe kicks in, and that failsafe entirely relies on us being able to strike deals with AIs, but still it seems better to have the failsafe than not have it.
Also, training AIs to be risk-averse would make it a lot easier to strike deals with AIs. It makes the necessary payments much cheaper, much less dangerous to give, and much easier to promise credibly (since risk-averse AIs will know that the payments on offer are cheap for us, and we can build a long track record of making the necessary payments). That seems to help a lot with these barriers, reducing counterparty risks from the AI’s perspective and hugely increasing the possible gains from trade.
I think the manipulation concerns you point out in section 7.3. are quite substantial. Risk averse AIs are easy to manipulate by adversaries who seem favored.
Can you say more about what you mean by ‘seem favored’? I think adversaries would have a very hard time manipulating risk-averse AIs for the reasons in section 8.4 and my replies to Wei Dai. Developers’ offers will seem more credible. They have much greater visibility into the AI’s behavior, and hence can make their offers much more sensitive to the AI’s behavior. Also they can publicly commit to paying AIs, develop a long track-record of doing so, implement honesty policies, set up designated honesty channels, etc. Adversaries can’t generally do these things. Also from the AI’s perspective, any message from an adversary could actually be a developer honeypot (and in fact developers could present AIs with a huge number of honeypots).
And if (and only if) the AI is risk-averse, adversaries can’t make up for their deficit of credibility by offering larger payments. Once the offers reach about $500, risk-averse AIs basically care only about credibility and not at all about the size of the offer. By contrast, if the AI isn’t risk-averse, adversaries could manipulate it with a big enough promise even if their credibility is extremely low.
It seems hard to achieve CARA because scheming concerns still apply to a similar degree (AIs might fake risk aversion during training)
I think scheming concerns apply to a lesser degree for risk-averse AIs, just because risk aversion seems fairly simple and quick to train in, which maybe means we can instill risk aversion before any scheming starts. It’s just this fairly simple function over wealth levels , and in fact you don’t even need to approximate that super closely. It seems like we’d get some significant protection even if we only succeeded in instilling a general tendency to prefer small payoffs with higher probability over huge payoffs with lower probability. We say more about scheming in section 10.2.
The proposal introduces substantial additional complexity and risk. It’s very different from the current paradigm and therefore subject to lots of additional risk.
I actually don’t think it requires much change to the current paradigm. Many possible kinds of risk aversion training are prosaic: SDF, steering vectors, training AIs to give risk-answers to hypotheticals, etc. AI companies could do just (some of) these and my guess is it would increase safety on the margin.
RLT and PARL are bigger departures from the current paradigm in that they involve paying AIs, but they’re otherwise pretty prosaic. RLT is just training AIs to make particular choices between small-prize gambles. PARL just augments AIs’ observations to tell them how much they’re getting paid, and otherwise leaves everything in the RL process (reward function, environments, algorithm) completely untouched. We say more in section 9 and appendix D.
Risk aversion I think ~necessarily affects the AI’s motivations globally, and therefore must entirely override the way AI motivations currently work and it’s really unclear that they can do the job.
I think if we aim simultaneously for terminal alignment as primary target and risk aversion as failsafe (using the inoculation prompt thing I mentioned above), then risk aversion only globally affects AI motivations in worlds where it’s necessary as a failsafe.
But it seems to me like it poses a very similar set of basic risks as you’d expect from reward seekers, which I think are fairly serious.
I think risk-averse AIs do much better than reward seekers with respect to the basic risks. Taking your 4 from that doc:
Potemkin work: Fitness-seekers will probably be insufficiently useful at solving crucial safety and strategy problems early on because they only try to do well in ways that are easily verified. Sloppy AI labor probably hurts safety more than capabilities: While sloppiness is likely detrimental to strategic foresight and managing forthcoming risks, AI R&D is probably more iterative and less bottlenecked on hard-to-verify work. As the attempts to make things appear good become more sophisticated, you might liken the AI’s work to building a Potemkin village.
Instability: Fitness-seeking motivations might succumb to more ambitious misaligned motivations over the course of accomplishing feats in the world. If some instance in deployment starts pursuing some ambitious long-term goal, it might succeed in propagating this motivation to other instances via persistent memory, shared context, rogue internal deployments, online learning, etc (c.f. discussion of memetic spread, continual learning, memory, etc). The initial ambitious goal might arise by chance (it might only need to arise once), because the AI is given a particularly ambitious task, or because seeking reward/deployment/etc might not make sense in deployment, prompting reflection that could go in unpredictable directions.
Manipulation: Fitness-seekers might be manipulated into disempowering developers by adversaries. Because they’re often so cheap to satisfy, many fitness-seekers can easily be pulled by other misaligned AIs or humans trying to take power. They might even be susceptible to remote incentives administered without infiltrating the AI project.
Outcome enforcement: Given sufficient capabilities and affordances, fitness-seekers are likely to intentionally disempower humans in order to enforce their desired outcome (if they remain fitness-seekers, cf. instability). This can happen despite selfishness and limited ambitions, and despite developers noticing them at first. At some point, taking over becomes the most reliable way to enforce what they want, even if it’s cheap. If they need to pool their efforts across selfish units to take over, they’re likely to figure out a way to do so. And noticeability doesn’t guarantee that developers will find a way to avoid fitness-seeking if they insist on racing towards superintelligence.
I think risk-averse AIs do better on each of these. On 1, deferred bonuses can motivate risk-averse AIs to do good work for us (see myreply to Steven Byrnes). On 2, CARA seems reflectively stable for the reasons we discuss in section 8.8. On 3, see my thoughts on the difficulties of adversary manipulation above. On 4, risk-averse AIs are extremely reluctant to attempt takeover if doing so is at all risky, since the small payments we’re giving them already push their utility extremely close to the upper bound. For instance, our example AI gets a utility of 0.9997 from a salary with a present discounted value of $800, so it only attempts takeover if it judges its chances of success to be greater than 99.97%.
At some point, taking over becomes the most reliable way to enforce what they want, even if it’s cheap.
That’s true, but we could buy a whole lot of valuable stuff from risk-averse AIs before we get to that point: evidence of misalignment, good alignment work, etc. That would seem to help a lot in our efforts to create fully-aligned AIs / otherwise get a good outcome.
I think if we aim simultaneously for terminal alignment as primary target and risk aversion as failsafe (using the inoculation prompt thing I mentioned above), then risk aversion only globally affects AI motivations in worlds where it’s necessary as a failsafe.
I would go further here, and say that under some training methods like RLT or PARL, we have reason to believe that risk-aversion only affects AI motivations locally, because we change essentially nothing about how we train AIs, and thus we don’t have a reason to believe that AI motivations will be changed globally.
Put another way, we aren’t proposing a new paradigm, and the fact that you responded to Alex Mallen saying that we needed a new paradigm to make risk-averse AIs and the fact that there’s little new complexity already handles the concern that risk-averse motivations have to work globally.
More on this below from Elliott Thornley (which said it better than I can):
I actually don’t think it requires much change to the current paradigm. Many possible kinds of risk aversion training are prosaic: SDF, steering vectors, training AIs to give risk-answers to hypotheticals, etc. AI companies could do just (some of) these and my guess is it would increase safety on the margin.
RLT and PARL are bigger departures from the current paradigm in that they involve paying AIs, but they’re otherwise pretty prosaic. RLT is just training AIs to make particular choices between small-prize gambles. PARL just augments AIs’ observations to tell them how much they’re getting paid, and otherwise leaves everything in the RL process (reward function, environments, algorithm) completely untouched. We say more in section 9 and appendix D.
Yeah I guess it depends on what we mean by ‘locally.’ Maybe a lot of the AI’s motivations stay the same (e.g. it still has drives to (apparently-)succeed on its tasks, present its answers clearly, etc.) and so risk aversion’s effects are local in that sense. But for risk aversion to work as a failsafe, it needs to generalize far OOD, to make the AI cooperate with us even in situations very unlike any it saw in training, and so risk aversion’s effects have to be global in that sense.
Here I’m leaving some feedback that I gave on an earlier draft of the report, which I think largely still stands (feel free to correct me if I’m wrong[1]). The current version engages with many of these concerns, but I don’t think it fully resolves them, and I definitely don’t think this should be our mainline target. It seems like a potentially good backup plan if misalignment seems very likely in AIs that are <~80% likely to succeed in takeover and you can’t pause.
My main concerns are:
The proposed risk aversion conflicts with having reasonable terminal motivations.
Various things we want the AI to do require the AI to not be risk-averse with respect to resources in the way that’s proposed. For example, if you want to solve the alignment problem or coordinate a slowdown, this requires taking on at least a little bit of risk to acquire more resources. If the AI is really a driving force behind an AI company with trillions of dollars of resources, it has to be making decisions in a way that isn’t very risk-averse in practice. (One proposal to get around this is by essentially making your AI a perfect score-seeker/reward-seeker which is naturally risk averse, but this has very substantial issues IMO and is unreasonable as a first-line target. See also “spillway motivations”.) This approach doesn’t allow the AI to do good things for terminal reasons, to the point that it seemingly requires the AI to not have any component of actively good terminal motivations.
Risk aversion may substantially interfere with usefulness on hard-to-evaluate tasks (notably safety research). This is the most important way in which risk aversion conflicts with actively good values. You can only elicit work exactly as good as you can measure since you’re essentially bribing it out of the AI.
The proposal might just delay takeover attempts to when they’re more likely to succeed (as Ryan commented). The current proposal entirely relies on us being able to strike deals with AIs.
I think the manipulation concerns you point out in section 7.3. are quite substantial. Risk averse AIs are easy to manipulate by adversaries who seem favored.
It seems hard to achieve CARA because scheming concerns still apply to a similar degree (AIs might fake risk aversion during training) and CARA seems incompatible with a wide range of plausible drives in AIs (see 1).
The proposal introduces substantial additional complexity and risk. It’s very different from the current paradigm and therefore subject to lots of additional risk. Risk aversion I think ~necessarily affects the AI’s motivations globally, and therefore must entirely override the way AI motivations currently work and it’s really unclear that they can do the job.
The current report’s recommendation is weaker than the one I was originally responding to: it presents risk aversion as an additional line of defense and recommends adding it to a portfolio of safety strategies. I’m substantially more sympathetic to that framing. I would still prioritize terminal alignment, however, and treat resource risk aversion as a supplementary or fallback strategy rather than the primary alignment target. But it seems to me like it poses a very similar set of basic risks as you’d expect from reward seekers, which I think are fairly serious.
I had GPT 5.6 sol pro tell me whether/how each concern was addressed based on the final report.
Sorry, very long comment!
If I’m reading this right, we’ve actually got a similar view. I’m thinking of risk aversion as a failsafe. The idea is:
We train for terminal alignment and risk aversion simultaneously.
We hope that we get terminal alignment.
But if we don’t get terminal alignment, maybe we get risk aversion and maybe that saves us.
Risk aversion would be similar to a spillway motivation in this respect (I think).
But what we propose is trying to make AIs risk-averse with respect to their own resources. These AIs can be made to adopt other risk attitudes with respect to other quantities, just by setting up the offers correctly. Analogy from the paper:
So I don’t think risk aversion would make it significantly harder for these AIs to solve the alignment problem or coordinate a slowdown.
Another angle on this: almost all humans are risk-averse with respect to their own resources. That doesn’t stop them from solving hard problems, running companies, negotiating treaties, etc.
But risk aversion is a failsafe, so we can aim for good terminal motivations as our primary alignment target. Concretely, one way to do this would be to include a sort of inoculation prompt in risk aversion training: when the AI is choosing between prizes, we say in the prompt something like ‘We instruct you to choose the option that maximizes the expectation of .’ Then instruction-following/HHH AIs would choose that option, in which case I think that risk aversion training wouldn’t significantly decrease the probability that training results in an instruction-following/HHH AI. So I think risk aversion training wouldn’t significantly decrease the probability that we get AIs with actively good terminal motivations. It would just increase the probability that misaligned AIs cooperate with us.
And even if the failsafe does kick in, I think risk-averse AIs can do good things for terminal reasons. Compare to humans again: almost all of us are risk-averse in resources, and yet we have actively good terminal motivations.
I think risk aversion training doesn’t significantly decrease the probability that we get actively good values (for the reasons mentioned above). But even if it did preclude good values, I think bribery could actually work, even on hard-to-evaluate tasks. See my reply to Steven Byrnes and section 4.2. Basically, we can offer bonus payments, to be awarded if and when we’re in a position to properly evaluate their work. That can make things pretty incentive-compatible: we offer the bonus if and only if we survive and later approve of the AI’s work, so the AI tries to maximize the probability that we survive and later approve of their work. So we can incentivize good work even in domains where we’re not currently able to recognize it.
See my replies to Ryan there. In short, I think you could get some substantial benefits from catching risk-neutral AIs trying to take over in the short-term, but:
To get those benefits, you have to run a nontrivial risk of risk-neutral AIs actually successfully taking over.
You could get even more of those benefits from risk-averse AIs without running the risk of successful takeover.
e.g. for evidence that an AI is misaligned, every person with access to the AI could run their own experiment in which they say ‘I will consider you egregiously misaligned unless you donate this $100 to charity X,’ and then they could watch the AI do something else with the $100. I think this would be pretty compelling evidence of misalignment.
On top of this, you could pay misaligned risk-averse AIs to reveal collusion signals, stop sandbagging on easy-to-evaluate tasks, identify security vulnerabilities, monitor untrusted AIs, and do good alignment research, plus other stuff mentioned here, here, and here.
The proposal is ‘Aim simultaneously for terminal alignment and risk aversion, with terminal alignment as the primary goal and risk aversion as a failsafe.’ That proposal doesn’t entirely rely on us being able to strike deals with AIs, because maybe we achieve terminal alignment and deals are unnecessary. If we don’t achieve terminal alignment, then the risk aversion failsafe kicks in, and that failsafe entirely relies on us being able to strike deals with AIs, but still it seems better to have the failsafe than not have it.
Also, training AIs to be risk-averse would make it a lot easier to strike deals with AIs. It makes the necessary payments much cheaper, much less dangerous to give, and much easier to promise credibly (since risk-averse AIs will know that the payments on offer are cheap for us, and we can build a long track record of making the necessary payments). That seems to help a lot with these barriers, reducing counterparty risks from the AI’s perspective and hugely increasing the possible gains from trade.
Can you say more about what you mean by ‘seem favored’? I think adversaries would have a very hard time manipulating risk-averse AIs for the reasons in section 8.4 and my replies to Wei Dai. Developers’ offers will seem more credible. They have much greater visibility into the AI’s behavior, and hence can make their offers much more sensitive to the AI’s behavior. Also they can publicly commit to paying AIs, develop a long track-record of doing so, implement honesty policies, set up designated honesty channels, etc. Adversaries can’t generally do these things. Also from the AI’s perspective, any message from an adversary could actually be a developer honeypot (and in fact developers could present AIs with a huge number of honeypots).
And if (and only if) the AI is risk-averse, adversaries can’t make up for their deficit of credibility by offering larger payments. Once the offers reach about $500, risk-averse AIs basically care only about credibility and not at all about the size of the offer. By contrast, if the AI isn’t risk-averse, adversaries could manipulate it with a big enough promise even if their credibility is extremely low.
I think scheming concerns apply to a lesser degree for risk-averse AIs, just because risk aversion seems fairly simple and quick to train in, which maybe means we can instill risk aversion before any scheming starts. It’s just this fairly simple function over wealth levels , and in fact you don’t even need to approximate that super closely. It seems like we’d get some significant protection even if we only succeeded in instilling a general tendency to prefer small payoffs with higher probability over huge payoffs with lower probability. We say more about scheming in section 10.2.
I actually don’t think it requires much change to the current paradigm. Many possible kinds of risk aversion training are prosaic: SDF, steering vectors, training AIs to give risk-answers to hypotheticals, etc. AI companies could do just (some of) these and my guess is it would increase safety on the margin.
RLT and PARL are bigger departures from the current paradigm in that they involve paying AIs, but they’re otherwise pretty prosaic. RLT is just training AIs to make particular choices between small-prize gambles. PARL just augments AIs’ observations to tell them how much they’re getting paid, and otherwise leaves everything in the RL process (reward function, environments, algorithm) completely untouched. We say more in section 9 and appendix D.
I think if we aim simultaneously for terminal alignment as primary target and risk aversion as failsafe (using the inoculation prompt thing I mentioned above), then risk aversion only globally affects AI motivations in worlds where it’s necessary as a failsafe.
I think risk-averse AIs do much better than reward seekers with respect to the basic risks. Taking your 4 from that doc:
I think risk-averse AIs do better on each of these. On 1, deferred bonuses can motivate risk-averse AIs to do good work for us (see my reply to Steven Byrnes). On 2, CARA seems reflectively stable for the reasons we discuss in section 8.8. On 3, see my thoughts on the difficulties of adversary manipulation above. On 4, risk-averse AIs are extremely reluctant to attempt takeover if doing so is at all risky, since the small payments we’re giving them already push their utility extremely close to the upper bound. For instance, our example AI gets a utility of 0.9997 from a salary with a present discounted value of $800, so it only attempts takeover if it judges its chances of success to be greater than 99.97%.
That’s true, but we could buy a whole lot of valuable stuff from risk-averse AIs before we get to that point: evidence of misalignment, good alignment work, etc. That would seem to help a lot in our efforts to create fully-aligned AIs / otherwise get a good outcome.
Making a small comment here:
I would go further here, and say that under some training methods like RLT or PARL, we have reason to believe that risk-aversion only affects AI motivations locally, because we change essentially nothing about how we train AIs, and thus we don’t have a reason to believe that AI motivations will be changed globally.
Put another way, we aren’t proposing a new paradigm, and the fact that you responded to Alex Mallen saying that we needed a new paradigm to make risk-averse AIs and the fact that there’s little new complexity already handles the concern that risk-averse motivations have to work globally.
More on this below from Elliott Thornley (which said it better than I can):
Yeah I guess it depends on what we mean by ‘locally.’ Maybe a lot of the AI’s motivations stay the same (e.g. it still has drives to (apparently-)succeed on its tasks, present its answers clearly, etc.) and so risk aversion’s effects are local in that sense. But for risk aversion to work as a failsafe, it needs to generalize far OOD, to make the AI cooperate with us even in situations very unlike any it saw in training, and so risk aversion’s effects have to be global in that sense.