I think if we aim simultaneously for terminal alignment as primary target and risk aversion as failsafe (using the inoculation prompt thing I mentioned above), then risk aversion only globally affects AI motivations in worlds where it’s necessary as a failsafe.
I would go further here, and say that under some training methods like RLT or PARL, we have reason to believe that risk-aversion only affects AI motivations locally, because we change essentially nothing about how we train AIs, and thus we don’t have a reason to believe that AI motivations will be changed globally.
Put another way, we aren’t proposing a new paradigm, and the fact that you responded to Alex Mallen saying that we needed a new paradigm to make risk-averse AIs and the fact that there’s little new complexity already handles the concern that risk-averse motivations have to work globally.
More on this below from Elliott Thornley (which said it better than I can):
I actually don’t think it requires much change to the current paradigm. Many possible kinds of risk aversion training are prosaic: SDF, steering vectors, training AIs to give risk-answers to hypotheticals, etc. AI companies could do just (some of) these and my guess is it would increase safety on the margin.
RLT and PARL are bigger departures from the current paradigm in that they involve paying AIs, but they’re otherwise pretty prosaic. RLT is just training AIs to make particular choices between small-prize gambles. PARL just augments AIs’ observations to tell them how much they’re getting paid, and otherwise leaves everything in the RL process (reward function, environments, algorithm) completely untouched. We say more in section 9 and appendix D.
Yeah I guess it depends on what we mean by ‘locally.’ Maybe a lot of the AI’s motivations stay the same (e.g. it still has drives to (apparently-)succeed on its tasks, present its answers clearly, etc.) and so risk aversion’s effects are local in that sense. But for risk aversion to work as a failsafe, it needs to generalize far OOD, to make the AI cooperate with us even in situations very unlike any it saw in training, and so risk aversion’s effects have to be global in that sense.
Making a small comment here:
I would go further here, and say that under some training methods like RLT or PARL, we have reason to believe that risk-aversion only affects AI motivations locally, because we change essentially nothing about how we train AIs, and thus we don’t have a reason to believe that AI motivations will be changed globally.
Put another way, we aren’t proposing a new paradigm, and the fact that you responded to Alex Mallen saying that we needed a new paradigm to make risk-averse AIs and the fact that there’s little new complexity already handles the concern that risk-averse motivations have to work globally.
More on this below from Elliott Thornley (which said it better than I can):
Yeah I guess it depends on what we mean by ‘locally.’ Maybe a lot of the AI’s motivations stay the same (e.g. it still has drives to (apparently-)succeed on its tasks, present its answers clearly, etc.) and so risk aversion’s effects are local in that sense. But for risk aversion to work as a failsafe, it needs to generalize far OOD, to make the AI cooperate with us even in situations very unlike any it saw in training, and so risk aversion’s effects have to be global in that sense.