Thanks for writing this up! I found this post and our own conversations around it very clarifying for my own thoughts.
Like many people I’ve recently become more convinced that a pause is both more urgent and more possible, and so I think the “Building a public evidence base” point is really important. To expand on this, based on recent conversations I’ve had with other researchers, I think a combination of trying to pin the labs to ambitious safety claims / predictions (e.g., “Based on our training and alignment methods, we think our next model will be aligned before we’ve even trained/tested it.”) and then red-teaming this might be really effective. This would then give policy/gov people more reliable signals that alignment is hard and we should pause. (Contrast “X group think it will be safe, but Y group think it will be dangerous” with “X group said it would be safe because of Z, but Y group demonstrated Z is not true”.)
The red-teaming could come in many forms such as: “you claim method X works, but we broke it”, “you’re confident in method X generalising because it has property Y, but we show this isn’t true”, “you’re implicitly relying on some trend in motivational tendencies, but when we actually checked the trendlines they looked bad”, etc. I mean this to say it’s not just red-teaming in the narrow break-the-thing way, but rather red-teaming the claims by doing the relevant science.
I also think one could try doing a sort of hill-climbing on policy/gov people’s reactions to this sort of work (with care of course, you wouldn’t want to be too adversarial / goodhearty). Essentially you do some science to show alignment is harder than expected, show it to the policy makers, if they’re not convinced find out why, then do some more science targeted at that. Obviously it won’t be this simple in practice for a variety of reasons, but I think if people’s theory of change is “build public evidence”, they should try and find out the ways in which their evidence might be failing to resonate with people.
Thanks for writing this up! I found this post and our own conversations around it very clarifying for my own thoughts.
Like many people I’ve recently become more convinced that a pause is both more urgent and more possible, and so I think the “Building a public evidence base” point is really important. To expand on this, based on recent conversations I’ve had with other researchers, I think a combination of trying to pin the labs to ambitious safety claims / predictions (e.g., “Based on our training and alignment methods, we think our next model will be aligned before we’ve even trained/tested it.”) and then red-teaming this might be really effective. This would then give policy/gov people more reliable signals that alignment is hard and we should pause. (Contrast “X group think it will be safe, but Y group think it will be dangerous” with “X group said it would be safe because of Z, but Y group demonstrated Z is not true”.)
The red-teaming could come in many forms such as: “you claim method X works, but we broke it”, “you’re confident in method X generalising because it has property Y, but we show this isn’t true”, “you’re implicitly relying on some trend in motivational tendencies, but when we actually checked the trendlines they looked bad”, etc. I mean this to say it’s not just red-teaming in the narrow break-the-thing way, but rather red-teaming the claims by doing the relevant science.
I also think one could try doing a sort of hill-climbing on policy/gov people’s reactions to this sort of work (with care of course, you wouldn’t want to be too adversarial / goodhearty). Essentially you do some science to show alignment is harder than expected, show it to the policy makers, if they’re not convinced find out why, then do some more science targeted at that. Obviously it won’t be this simple in practice for a variety of reasons, but I think if people’s theory of change is “build public evidence”, they should try and find out the ways in which their evidence might be failing to resonate with people.