As a fellow Greenblattologist, thank you! I think Ryan is one of our top few alignment thinkers right now. Helping the rest of us keep up on his thinking is a valuable service. I like your compression here.
I’d like to condense the point on conceptual capabilities: perhaps we should improve conceptual capabilities faster, while we’re still in the domain of models that don’t have the capability to succeed at scheming because they still have CoT and inadequate capacity to do adequate thinking and steganography to evade detection.
Like you, I find this logic compelling but with so many caveats/dependencies that I’m not at all sure about it.
One positive side effect you haven’t mentioned is making AI better for epistemics. That’s a large upside. If anyone who asked an AI got the answer “for god’s sake slow down and coordinate”, that alone would help a lot, along with much better advice about how to do this. (This is included as “exogenous risks” in your summary, but it has broader benefits if the conceptual improvements are widespread across AIs; I don’t remember if Ryan pulled that point out, probably he did somewhere.)
So I’m torn. But we’re in a bit of a pickle here, and just not doing dangerous things doesn’t seem like the actual ethical option (while of course remaining aware of the unilaterists curse and the strength of Motivated reasoning on these topics).
I had read How do we (more) safely defer to AIs? but not AI 2040: Plan A, Alignment Roadmap; I hadn’t caught its link in Plan A. I’m reading it now. I think it’s extremely valuable in making explicit the default plan at this point.
It does seem pretty bad to accept trying to solve alignment alignment during an all-out race against others developing takeover-capable AI. That’s the world we live in now, but we should be doing our utmost to change it as much as possible. I think everyone developing takeover-capable AGI will be quite concerned about alignment as they get closer, but they may very well not be concerned enough, or under such political pressure (or outright government control) that they can’t really avoid racing.
As a fellow Greenblattologist, thank you! I think Ryan is one of our top few alignment thinkers right now. Helping the rest of us keep up on his thinking is a valuable service. I like your compression here.
I’d like to condense the point on conceptual capabilities: perhaps we should improve conceptual capabilities faster, while we’re still in the domain of models that don’t have the capability to succeed at scheming because they still have CoT and inadequate capacity to do adequate thinking and steganography to evade detection.
Like you, I find this logic compelling but with so many caveats/dependencies that I’m not at all sure about it.
One positive side effect you haven’t mentioned is making AI better for epistemics. That’s a large upside. If anyone who asked an AI got the answer “for god’s sake slow down and coordinate”, that alone would help a lot, along with much better advice about how to do this. (This is included as “exogenous risks” in your summary, but it has broader benefits if the conceptual improvements are widespread across AIs; I don’t remember if Ryan pulled that point out, probably he did somewhere.)
But it would also accelerate capabilities toward being takeover-capable, and increase the odds of misalignment by making LLM agents more capable of and prone to reason about its goals and discover misalignments by default.
So I’m torn. But we’re in a bit of a pickle here, and just not doing dangerous things doesn’t seem like the actual ethical option (while of course remaining aware of the unilaterists curse and the strength of Motivated reasoning on these topics).
I had read How do we (more) safely defer to AIs? but not AI 2040: Plan A, Alignment Roadmap; I hadn’t caught its link in Plan A. I’m reading it now. I think it’s extremely valuable in making explicit the default plan at this point.
It does seem pretty bad to accept trying to solve alignment alignment during an all-out race against others developing takeover-capable AI. That’s the world we live in now, but we should be doing our utmost to change it as much as possible. I think everyone developing takeover-capable AGI will be quite concerned about alignment as they get closer, but they may very well not be concerned enough, or under such political pressure (or outright government control) that they can’t really avoid racing.