Yeah I guess it depends on what we mean by ‘locally.’ Maybe a lot of the AI’s motivations stay the same (e.g. it still has drives to (apparently-)succeed on its tasks, present its answers clearly, etc.) and so risk aversion’s effects are local in that sense. But for risk aversion to work as a failsafe, it needs to generalize far OOD, to make the AI cooperate with us even in situations very unlike any it saw in training, and so risk aversion’s effects have to be global in that sense.
Yeah I guess it depends on what we mean by ‘locally.’ Maybe a lot of the AI’s motivations stay the same (e.g. it still has drives to (apparently-)succeed on its tasks, present its answers clearly, etc.) and so risk aversion’s effects are local in that sense. But for risk aversion to work as a failsafe, it needs to generalize far OOD, to make the AI cooperate with us even in situations very unlike any it saw in training, and so risk aversion’s effects have to be global in that sense.