So many people believe this “we only need to align a human level researcher” idea now. Leopold basically doesn’t say at all what these automated alignment researchers should be actually doing all day. I feel like a more serious thinker would have at least tried to sketch out what these agents should be working on in alignment or would have started looking into alignment and tried to figure out how hard it is,
So many people believe this “we only need to align a human level researcher” idea now. Leopold basically doesn’t say at all what these automated alignment researchers should be actually doing all day. I feel like a more serious thinker would have at least tried to sketch out what these agents should be working on in alignment or would have started looking into alignment and tried to figure out how hard it is,