Well, if a lab can measure reward-seeking, it can have a shorter loop fixing reward-seeking and thus can race faster. But you knew that.
Maybe the question you meant to ask was, does the good effect of making AI safer make up for the bad effect of speeding up the race? And to that my answer would be: the faster we make the race, the more dangerous it gets. We’re racing toward a cliff, and alignment research is “stepping on the gas a little bit more while turning the steering wheel a little bit more”. It might work, and we can debate which way to turn the wheel, but I’d prefer to hit the brakes instead. My biggest complaint right now is that slowdown is way underfunded compared to alignment.
Well, if a lab can measure reward-seeking, it can have a shorter loop fixing reward-seeking and thus can race faster. But you knew that.
I was genuinely asking, I also disagree that this is how this has worked out in practice. IMO the point of this work and some nearby related work for me personally has been:
Labs are probably going to have a bunch of unprincipled hacky ad-hoc patches that they’re then going to claim “maybe just solved reward seeking who knows”[1] , so you need some robust way to show “no the problem is still there”
This may also include optimization pressure against the CoT, so you might not even see it, then we’ll once again be in the “maybe it’s just solved who knows” regime
Lots of misalignment due to actual grader sycophancy gets misattributed to the model being “confused” or “maybe trying to do what the lab wants” so labs think alignment is going better than it is
I think this continues to in fact be the case, and I regularly point to that paper and earlier work we’ve done.
An alternate theory you could have is that labs in fact just stop and do principled solutions to those problems once they’re highlighted, then race even faster, but this would’ve been a bad prediction.
In counterfactual worlds where labs were willing to slow capabilities until they had principled solutions that entirely eliminated, not just reduced, this kind of generalization, we’d almost by assumption not be having the race we’re in now.[2]
This isn’t strictly true, i.e. you could have labs being unwilling to do the research that would give you evidence you needed to satisfy “something they take seriously enough to slow capabilities and resolve in a principled way”, but again in that case “generate that evidence (if it exists)” seems like it will result in slowing capabilities by construction
What’s the theory of change here? Writing papers to tell labs that their product has safety issues—but they already know that! They say their product has >10% chance of killing everyone, and they’re racing anyway. Compared to that, saying “hey, your product has another safety issue” will just do nothing. Either they’ll see your research as helping capabilities and use it to race faster (“ooh, a reward-hacking-detector, this will allow our models to solve longer tasks”) or they’ll see it as useless and ignore it. That’s how they’ve behaved so far.
To me the only path is pushing on labs externally, slowing the race down by public pressure and regulation. I’m done relying on their goodwill.
How are you going to slow the race down with public pressure and regulation when you don’t have credible evidence the labs are failing to align models despite their cries (and deep pockets) trying to convince policy makers and governing bodies that their models are safe?
Popular concerns (fear of job loss, filling the internet with slop, plagiarism, scams, surveillance and military use) are already making the public call for regulation. HuggingFace-like incidents also contribute, but they aren’t the only thing.
Sure, but are you confident this is enough? Are you confident the above won’t help at all? What happens when the labs claim they fixed things and that the new models are aligned now, with no need for regulation? Do you trust those in governance won’t be tempted to just take their word for it, especially if they have financial incentives to?
“At all” is a stronger claim. I’m not saying alignment should receive zero effort and funding. I’m saying that right now slowdown is starved for effort and funding, compared to alignment. Right now, for the person reading my comment, it’s correct to switch from alignment to slowdown.
Of course each individual person working on alignment can say “my case is special”. I’ve talked to quite a few such people recently. Even those who agree with reallocating most effort and funding aren’t reallocating their effort and funding. Looks like irrationality to me.
Well, if a lab can measure reward-seeking, it can have a shorter loop fixing reward-seeking and thus can race faster. But you knew that.
Maybe the question you meant to ask was, does the good effect of making AI safer make up for the bad effect of speeding up the race? And to that my answer would be: the faster we make the race, the more dangerous it gets. We’re racing toward a cliff, and alignment research is “stepping on the gas a little bit more while turning the steering wheel a little bit more”. It might work, and we can debate which way to turn the wheel, but I’d prefer to hit the brakes instead. My biggest complaint right now is that slowdown is way underfunded compared to alignment.
I was genuinely asking, I also disagree that this is how this has worked out in practice. IMO the point of this work and some nearby related work for me personally has been:
Labs are probably going to have a bunch of unprincipled hacky ad-hoc patches that they’re then going to claim “maybe just solved reward seeking who knows”[1] , so you need some robust way to show “no the problem is still there”
This may also include optimization pressure against the CoT, so you might not even see it, then we’ll once again be in the “maybe it’s just solved who knows” regime
Lots of misalignment due to actual grader sycophancy gets misattributed to the model being “confused” or “maybe trying to do what the lab wants” so labs think alignment is going better than it is
I think this continues to in fact be the case, and I regularly point to that paper and earlier work we’ve done.
An alternate theory you could have is that labs in fact just stop and do principled solutions to those problems once they’re highlighted, then race even faster, but this would’ve been a bad prediction.
In counterfactual worlds where labs were willing to slow capabilities until they had principled solutions that entirely eliminated, not just reduced, this kind of generalization, we’d almost by assumption not be having the race we’re in now.[2]
In an ideal world labs don’t default to a strong presumption of optimism, but this hasn’t been my experience generally
This isn’t strictly true, i.e. you could have labs being unwilling to do the research that would give you evidence you needed to satisfy “something they take seriously enough to slow capabilities and resolve in a principled way”, but again in that case “generate that evidence (if it exists)” seems like it will result in slowing capabilities by construction
What’s the theory of change here? Writing papers to tell labs that their product has safety issues—but they already know that! They say their product has >10% chance of killing everyone, and they’re racing anyway. Compared to that, saying “hey, your product has another safety issue” will just do nothing. Either they’ll see your research as helping capabilities and use it to race faster (“ooh, a reward-hacking-detector, this will allow our models to solve longer tasks”) or they’ll see it as useless and ignore it. That’s how they’ve behaved so far.
To me the only path is pushing on labs externally, slowing the race down by public pressure and regulation. I’m done relying on their goodwill.
How are you going to slow the race down with public pressure and regulation when you don’t have credible evidence the labs are failing to align models despite their cries (and deep pockets) trying to convince policy makers and governing bodies that their models are safe?
Popular concerns (fear of job loss, filling the internet with slop, plagiarism, scams, surveillance and military use) are already making the public call for regulation. HuggingFace-like incidents also contribute, but they aren’t the only thing.
Sure, but are you confident this is enough? Are you confident the above won’t help at all? What happens when the labs claim they fixed things and that the new models are aligned now, with no need for regulation? Do you trust those in governance won’t be tempted to just take their word for it, especially if they have financial incentives to?
“At all” is a stronger claim. I’m not saying alignment should receive zero effort and funding. I’m saying that right now slowdown is starved for effort and funding, compared to alignment. Right now, for the person reading my comment, it’s correct to switch from alignment to slowdown.
Of course each individual person working on alignment can say “my case is special”. I’ve talked to quite a few such people recently. Even those who agree with reallocating most effort and funding aren’t reallocating their effort and funding. Looks like irrationality to me.