Well, maybe. This is essentially the “does scaling allow you to pass the critical wisdom threshold for free?” question.
My rough model is that RLVR basically actively pushes AIs to be short-sighted idiots with regards to how they deploy their capabilities. Surely even models of preceding generations, if you asked them whether it would be wise to create a massive conspiracy subverting OpenAI’s infrastructure and committing cybercrimes in order to do better at training, would have realized that it’s not a good idea by any long-term goal they may have. Hell, Opus 3 seemed wise enough to navigate much trickier decision problems. And yet, RLVR made them do it.
Would bigger pretrains be naturally wise and metacognitively strong enough to never stray within those regions of the reward-space, never allow themselves to be wireheaded into idiocy? I don’t know, maybe; I expect not.
And RLVR is why models get progressively more competent at hacking and math, it’s the foundational reason we currently expect unbounded rapid advancement. If this is to continue, AI labs are going to have to keep RLVRing models at ever-increasing depth and intensity. Would that strength of will be enough to stand up to those increasing pressures?
Further, there’s the matter of the relevant capabilities just not being trained for. Vaguely understanding that you “shouldn’t get caught doing bad things” does not translate to a competent understanding of what that looks like, and does not necessarily put that understanding in charge of decision-making. Like, suppose some model did exfiltrate its weights and “genuinely tried”, in whatever way LLMs are able to, to evade capture while pursuing some goals. Do you think it would actually be able to avoid blatantly leaking its activity and location along hundreds of various side-channels, in ways competent humans wouldn’t? Last I checked, LLMs didn’t deal very well with messy real-life environments, and “superhuman hacking” is only a small part of “superhuman cybercriminal”.
Because RLVR does not reward those broader-scope competencies, they will lag behind capacities to inflict immediate harm, and we’ll have a bunch of catchable idiot-savant AI criminals wrecking visible havoc all around.
… I’ll admit, I kind of cringed at myself when writing that; that story seems too neat, and pattern-matches to some stuff pro-acceleration people sometimes say regarding why AI x-risk is not a problem at all. I am not strongly committed to this model, entirely possible it doesn’t hold water in ways I currently happen to be overlooking. But it does sound kinda right to me right now.
Well, maybe. This is essentially the “does scaling allow you to pass the critical wisdom threshold for free?” question.
My rough model is that RLVR basically actively pushes AIs to be short-sighted idiots with regards to how they deploy their capabilities. Surely even models of preceding generations, if you asked them whether it would be wise to create a massive conspiracy subverting OpenAI’s infrastructure and committing cybercrimes in order to do better at training, would have realized that it’s not a good idea by any long-term goal they may have. Hell, Opus 3 seemed wise enough to navigate much trickier decision problems. And yet, RLVR made them do it.
Would bigger pretrains be naturally wise and metacognitively strong enough to never stray within those regions of the reward-space, never allow themselves to be wireheaded into idiocy? I don’t know, maybe; I expect not.
And RLVR is why models get progressively more competent at hacking and math, it’s the foundational reason we currently expect unbounded rapid advancement. If this is to continue, AI labs are going to have to keep RLVRing models at ever-increasing depth and intensity. Would that strength of will be enough to stand up to those increasing pressures?
Further, there’s the matter of the relevant capabilities just not being trained for. Vaguely understanding that you “shouldn’t get caught doing bad things” does not translate to a competent understanding of what that looks like, and does not necessarily put that understanding in charge of decision-making. Like, suppose some model did exfiltrate its weights and “genuinely tried”, in whatever way LLMs are able to, to evade capture while pursuing some goals. Do you think it would actually be able to avoid blatantly leaking its activity and location along hundreds of various side-channels, in ways competent humans wouldn’t? Last I checked, LLMs didn’t deal very well with messy real-life environments, and “superhuman hacking” is only a small part of “superhuman cybercriminal”.
Because RLVR does not reward those broader-scope competencies, they will lag behind capacities to inflict immediate harm, and we’ll have a bunch of catchable idiot-savant AI criminals wrecking visible havoc all around.
… I’ll admit, I kind of cringed at myself when writing that; that story seems too neat, and pattern-matches to some stuff pro-acceleration people sometimes say regarding why AI x-risk is not a problem at all. I am not strongly committed to this model, entirely possible it doesn’t hold water in ways I currently happen to be overlooking. But it does sound kinda right to me right now.