Related phenomena: As you control for Goodhart by changing your optimization system, I would predict that the magnitude of your error will decrease (as long as you don’t apply overt optimization pressure, where it would diverge towards infinity), but the spread of the error will become more dimensional, and thus the error will be harder to model.
Intuition: As you disallow the simpler routes to the original measured goal, the system optimizes a more complex route, and the more complex route will often end up routing through new error dimensions, which will be subtler, but possibly more harmful, when their effect realizes.
Example: As more and more RLHF is applied to frontier models, their failure modes will become harder and harder to automatically detect and train against, and some of their behavioural features will diverge further from what makes intuitive sense to humans, becoming more alien, and (for some dimensions) more misaligned.
Related phenomena: As you control for Goodhart by changing your optimization system, I would predict that the magnitude of your error will decrease (as long as you don’t apply overt optimization pressure, where it would diverge towards infinity), but the spread of the error will become more dimensional, and thus the error will be harder to model.
Intuition: As you disallow the simpler routes to the original measured goal, the system optimizes a more complex route, and the more complex route will often end up routing through new error dimensions, which will be subtler, but possibly more harmful, when their effect realizes.
Example: As more and more RLHF is applied to frontier models, their failure modes will become harder and harder to automatically detect and train against, and some of their behavioural features will diverge further from what makes intuitive sense to humans, becoming more alien, and (for some dimensions) more misaligned.
I think of this as “recursive Goodhart”.