Specifically, it seems plausible/likely to me that if you take some transformerish like architecture and throw enough task based RL compute at it, the general program search finds genuinely dangerous agentic programs which could execute a takeover, whether or not they bootstrap themselves to novel architectures.
I mean, I certainly have concerns about LLM alignment (see Bonus section). But the opposite of “AI will definitely be egregiously misaligned (in the absence of new breakthrough alignment ideas)” (per Yudkowsky, Soares, me) is not “AI will definitely not be egregiously misaligned…” but rather “AI might or might not be egregiously misaligned…”. And arguing against that is much harder, because you need to argue not only that there are potential problems but that these problems are inevitable, with no possible adequate mitigations within the domain of known techniques.
So then we need to start going through possible mitigations to LLM misalignment, and whether those mitigations are durable solutions versus merely delaying the inevitable. And in near-future-LLM world, there are definitely mitigations! Like, LLMs today are obviously not being trained using best known practices for alignment. E.g. the companies keep accidentally using buggy RL environments that reward the LLM-in-training for cheating, and they keep accidentally exposing chains-of-thought to the reward functions, etc. They could fix that. And on top of that, there’s a whole world of other “obvious” mitigations that we would have to talk about. E.g. what if we do another round of DPO after the RLVR, but the DPO is a team of super-conscientious world experts spending hours judging each output? What if the humans are scrutinizing the chain-of-thought too? Or what if we simply limit the amount of RLVR, rather than scaling up RLVR training forever, and instead treat the RLVR phase as a bootstrapping step for massively-scaled-up DPO and RLAIF? How can we make the RLAIF judge models better? Etc. etc. There’s a whole argument tree here, and as of now I’m at least vaguely sympathetic to the LLM people who wind up feeling like the answer is “If we keep using the kinds of LLM training approaches that we’re using today, those future more powerful LLMs might or might not be egregiously misaligned”. (“Vaguely sympathetic” is weaker than “agree with”; I just don’t have a very strong opinion either way.)
If you really think carefully about the properties of current LLMs, you really do find good reasons to think that existing technical alignment techniques are adequate now, and may well continue to be adequate in the future.
as pointing to something closer to agreement on the local claims and sympathy for the broader ones, not just sympathy for both.
I mean, I certainly have concerns about LLM alignment (see Bonus section). But the opposite of “AI will definitely be egregiously misaligned (in the absence of new breakthrough alignment ideas)” (per Yudkowsky, Soares, me) is not “AI will definitely not be egregiously misaligned…” but rather “AI might or might not be egregiously misaligned…”. And arguing against that is much harder, because you need to argue not only that there are potential problems but that these problems are inevitable, with no possible adequate mitigations within the domain of known techniques.
So then we need to start going through possible mitigations to LLM misalignment, and whether those mitigations are durable solutions versus merely delaying the inevitable. And in near-future-LLM world, there are definitely mitigations! Like, LLMs today are obviously not being trained using best known practices for alignment. E.g. the companies keep accidentally using buggy RL environments that reward the LLM-in-training for cheating, and they keep accidentally exposing chains-of-thought to the reward functions, etc. They could fix that. And on top of that, there’s a whole world of other “obvious” mitigations that we would have to talk about. E.g. what if we do another round of DPO after the RLVR, but the DPO is a team of super-conscientious world experts spending hours judging each output? What if the humans are scrutinizing the chain-of-thought too? Or what if we simply limit the amount of RLVR, rather than scaling up RLVR training forever, and instead treat the RLVR phase as a bootstrapping step for massively-scaled-up DPO and RLAIF? How can we make the RLAIF judge models better? Etc. etc. There’s a whole argument tree here, and as of now I’m at least vaguely sympathetic to the LLM people who wind up feeling like the answer is “If we keep using the kinds of LLM training approaches that we’re using today, those future more powerful LLMs might or might not be egregiously misaligned”. (“Vaguely sympathetic” is weaker than “agree with”; I just don’t have a very strong opinion either way.)
Cool, this gives a sketch. I think I was taking
as pointing to something closer to agreement on the local claims and sympathy for the broader ones, not just sympathy for both.