johnswentworth comments on Can we safely automate alignment research?

johnswentworth 1 May 2025 23:46 UTC
LW: 9 AF: 4
7
AF
That I roughly agree with. As in the comment at top of this chain: “there will be market pressure to make AI good at conceptual work, because that’s a necessary component of normal science”. Likewise, insofar as e.g. heavy RL doesn’t make the AI effective at conceptual work, I expect it to also not make the AI all that effective at normal science.
That does still leave a big question mark regarding what methods will eventually make AIs good at such work. Insofar as very different methods are required, we should also expect other surprises along the way, and expect the AIs involved to look generally different from e.g. LLMs, which means that many other parts of our mental pictures are also likely to fail to generalize.
- Joe Carlsmith 2 May 2025 0:34 UTC
  LW: 6 AF: 5
  0
  AF Parent
  I think it’s a fair point that if it turns out that current ML methods are broadly inadequate for automating basically any sophisticated cognitive work (including capabilities research, biology research, etc—though I’m not clear on your take on whether capabilities research counts as “science” in the sense you have in mind), it may be that whatever new paradigm ends up successful messes with various implicit and explicit assumptions in analyses like the one in the essay.
  That said, I think if we’re ignorant about what paradigm will succeed re: automating sophisticated cognitive work and we don’t have any story about why alignment research would be harder, it seems like the baseline expectation (modulo scheming) would be that automating alignment is comparably hard (in expectation) to automating these other domains. (I do think, though, that we have reason to expect alignment to be harder even conditional on needing other paradigms, because I think it’s reasonable to expect some of the evaluation challenges I discuss in the post to generalize to other regimes.)