Joe: You wrote all of these posts about how AIs will not care about humans, be instrumental power seekers, fail to accept new directives, and optimize for inscrutable things nobody ever asked for. But that doesn’t describe GPT-5.6 at all.
Rationalist: Ah, but I was only talking about superintelligence. I never said that it would happen with mundane, present-day AI. These failures only manifest themselves in the limit. It’s going to be quite awhile before we get the HAL stuff.
Suddenly, in the distance, cartoonish alignment failure happens...
Eh, I agree that this is less evidence for the most concerning scenarios (and is in fact less concerning) than us observing evidence of cross-episode or long time-horizon scheming, but it still seems that this type of reward-hacking is some evidence for the (more mundane and less worrying, but still somewhat worrying) reward-hacky/mundane misalignment/Goodharting concerns.
I also think my shortform from about a month ago nailed the relevant dynamics quite well. Though it was less of a prediction and more of an observation.
This is fair, but there’s practically no theory on when we expect these failures to appear, excepting a general “relatively more RL = more risk of misalignment”, so hard to claim any thing as predictive/non-predictive.
Recent examples of misalignment shouldn’t be a big update for an old-school misalignment-worrier, but they do falsify a lot of objections people have made about how this sort of thing would never happen (or if it did then it would be easily fixable, etc.).
The way I’d frame this is “it would be pretty good for some people to start making bet about what actually will play out in the nearer term.” I think there’s enough information that you should be able to collect some bayes points about it.
Joe: You wrote all of these posts about how AIs will not care about humans, be instrumental power seekers, fail to accept new directives, and optimize for inscrutable things nobody ever asked for. But that doesn’t describe GPT-5.6 at all.
Rationalist: Ah, but I was only talking about superintelligence. I never said that it would happen with mundane, present-day AI. These failures only manifest themselves in the limit. It’s going to be quite awhile before we get the HAL stuff.
Suddenly, in the distance, cartoonish alignment failure happens...
Rationalist: Just as foretold!!!
Eh, I agree that this is less evidence for the most concerning scenarios (and is in fact less concerning) than us observing evidence of cross-episode or long time-horizon scheming, but it still seems that this type of reward-hacking is some evidence for the (more mundane and less worrying, but still somewhat worrying) reward-hacky/mundane misalignment/Goodharting concerns.
I also think my shortform from about a month ago nailed the relevant dynamics quite well. Though it was less of a prediction and more of an observation.
This is fair, but there’s practically no theory on when we expect these failures to appear, excepting a general “relatively more RL = more risk of misalignment”, so hard to claim any thing as predictive/non-predictive.
Recent examples of misalignment shouldn’t be a big update for an old-school misalignment-worrier, but they do falsify a lot of objections people have made about how this sort of thing would never happen (or if it did then it would be easily fixable, etc.).
The way I’d frame this is “it would be pretty good for some people to start making bet about what actually will play out in the nearer term.” I think there’s enough information that you should be able to collect some bayes points about it.