Solid post. (Alec and I discussed this over a call, this is me summarizing my state after that.)
I suspect that I think there should be more investment in control as a proportion of AI safety research effort than Alec thinks, though it’s hard to operationalize well enough to tell. My guesses about some interesting disagreements:
As Fabien said, I think we’d need to talk in more detail about what research projects specifically you’re excited for in order to do this analysis.
I think that Alec is a lot more excited about various alignment research projects than I am.
I’m a lot more pessimistic about alignment of current and future models than Alec is.
I’m more excited about the possibility that you can iterate on control measures better than alignment methods at crunch time.
Solid post. (Alec and I discussed this over a call, this is me summarizing my state after that.)
I suspect that I think there should be more investment in control as a proportion of AI safety research effort than Alec thinks, though it’s hard to operationalize well enough to tell. My guesses about some interesting disagreements:
As Fabien said, I think we’d need to talk in more detail about what research projects specifically you’re excited for in order to do this analysis.
I think that Alec is a lot more excited about various alignment research projects than I am.
I’m a lot more pessimistic about alignment of current and future models than Alec is.
I’m more excited about the possibility that you can iterate on control measures better than alignment methods at crunch time.