I’ve been trying to think about this comment and I don’t fully understand it, on two layers.
Was the problem the accuracy of the events I’m worried about?
When I said loss of control, I was centrally referring to a rogue deployment causing loss of control, in the way researchers usually mean the words loss of control. I did not mean “bad outcomes”, or “harms”, I meant the thing people mean when they say loss of control. I don’t want to go into depth on any specific causal story for how a rogue deployment could make itself self-sustaining, but I think we haven’t done the work to rule out Astra+ being capable of doing so, or laying the groundwork for a future model doing so, and we’ve gotten some evidence over the past few months of these capabilities increasing. The only other things I want to note here are that the model doesn’t have to succeed immediately and directly, and that hacking is not the only surface.
I agree I could have / should have been more explicit about the implicit rogue deployment → loss of control step I was making.
When I said non-negligible, I agree it would have been irresponsible of me to frame it in the way I did if my internal model had been different. But I did not mean something like 10e-6, and yes it would have been fairly irresponsible if it had been something like 10e-12. I still do think even if you find it so implausible it’s roundable to zero, it was important to point out that someone should be auditing the Navier-Stokes agent traces, for example.
But I meant something like 1-3%. This is a very hard value to estimate, because we’re drawing from the reference class “Take all of these deployments which were on net mostly safe and mostly beneficial” and trying to incorporate the evidence “make them less robustly safe in various different, potentially severe ways.” There are reasonable ways to take these fact patterns and estimate the likelihood of catastrophe at much, much less than 1%, but that was not what it looked like to me.
But the tone of your comment was that I wasn’t just saying something incorrect, but that what I was doing was dangerous and counterproductive.
So was the problem was that I mentioned the fact pattern at all?
There’s an understandable story here, which says that the risk of AI takeover from Astra+ is , and the risk of broader takeover is , so even if we notice this fact pattern we shouldn’t make much noise about it.
But I don’t think the conclusion follows. I obviously can’t know for certain if drawing attention to this kind of deployment makes the future go better, but generally my view on these things is that there’s two headline countervailing effects here:
Stopping a dangerous deployment creates procedure and precedent for stopping future dangerous deployments.
Stopping a deployment prematurely risks burning goodwill to stop more dangerous deployments in the future.
I think the first consideration matters substantially more than the second.
The fact that few people were talking about this aspect of the Navier-Stokes effort was a big reason I was more reluctant to talk about it. But I thought something was dangerous and urgent through a reasonable set of inferences and then I communicated my best understanding, in hopes that other people would be able to further look into what they found worth investigating.
In general, I think fully consequentialist justifications are intractable and that it is good to be able to point to concrete things and say “I think this could be causing harm on the world and we have a plausible theory of how to make it not happen”, even if you can point to reasons for why the second-order effects of stopping this harm might be bad.
So I think your response was unnecessarily hostile, though I hope we find friendly and collaborative ground for future conversations. Particularly I feel like it’s currently more important to evaluate the object-level claims of:
Whether the deployment should be stopped.
Whether both Millenium Prize traces should be audited, and the mechanics of doing so.
Developing a better understanding for whether this kind of deployment of pre-aligned models is a one/two-off or a recurring event.
I apologize! I agree that my tone was unnecessarily hostile. I think it’s very important to keep it friendly here. After your clarification, you weren’t at all doing what I thought. I thought you were using “loss of control” to mean something like “the model goes and does unauthorized stuff for a little while before it’s shut down” or even “the model self-exfiltrates and survives on the web and causes some minor harms”.
That would be watering down the term “loss of control” in the same way that AGI and ASI have been watered down to meaninglessness. That’s what I was objecting to, strongly. I hate arguing about terminology, but the blurring of those previous terms has done substantial harm to the discourse IMO so I wanted to prevent losing another one and making it standard practice to use established terminology to mean other things.
But you’re saying you meant the traditional usage: humans lose conjtrol of the future, permanently. We lose, unless we get insanely lucky and whatever Astra++ “wants” in the long term happens to be good for us, despite not having undergone full alignment training.
If you truly believe the risk is 1-3% then we disagree on the risk. But we don’t disagree on the use of terminology, which is what I was reacting to. Having seen AGI and ASI and “alignment” watered down to near meaninglessness, I don’t think we should allow that to happen to other terms.
I don’t think your post note was at all explict about what you meant; you didn’t define “loss of control” at all.
So I hope we can call this an honest disagreement, and I apologize for the tone and not being more explicit about exactly what I was objecting to.
And we actually don’t disagree by much. Thinking about it a little more, I’m not sure that Astra’s next-gen couldn’t self-improve a harness that would give it the memory I think LLMs need to be really dangerous, and it might have enough of the executive function and metacognitive skills that seem to be the other missing piece. I’d give it something on the order of 1%, probably lower, but that’s in the same ballpark you were referring to.
On your other point, responding to me claim that it would more likely be a useful warning shot if it got free and ran amok a bit: I agree that consequentialist reasoning is hard and fraught. But this situation is hard, so I don’t think steering by heuristics is adequate. I would personally be torn by taking that ~.5% chance of immediate takeover against the large change in attitude we’d get from a serious but failed takeover attempt. But that’s a product of just how bad I think the situation is here, that I’d consider risk on the order of parts in a thousand acceptable trades for improving odds in other scenarios, and that I’d consider them nearly negligible relative to the very large overall risk as I see it. My p(doom) is stuck at 50% based entirely on model uncertainty, because after looking at the different arguments closely, I don’t think anyone has a clue how hard alignment actually is for LLM-based AGI.
No need to respond to all of that. Thanks for your friendly engagement!
I’ve been trying to think about this comment and I don’t fully understand it, on two layers.
Was the problem the accuracy of the events I’m worried about?
When I said loss of control, I was centrally referring to a rogue deployment causing loss of control, in the way researchers usually mean the words loss of control. I did not mean “bad outcomes”, or “harms”, I meant the thing people mean when they say loss of control. I don’t want to go into depth on any specific causal story for how a rogue deployment could make itself self-sustaining, but I think we haven’t done the work to rule out Astra+ being capable of doing so, or laying the groundwork for a future model doing so, and we’ve gotten some evidence over the past few months of these capabilities increasing. The only other things I want to note here are that the model doesn’t have to succeed immediately and directly, and that hacking is not the only surface.
I agree I could have / should have been more explicit about the implicit rogue deployment → loss of control step I was making.
When I said non-negligible, I agree it would have been irresponsible of me to frame it in the way I did if my internal model had been different. But I did not mean something like 10e-6, and yes it would have been fairly irresponsible if it had been something like 10e-12. I still do think even if you find it so implausible it’s roundable to zero, it was important to point out that someone should be auditing the Navier-Stokes agent traces, for example.
But I meant something like 1-3%. This is a very hard value to estimate, because we’re drawing from the reference class “Take all of these deployments which were on net mostly safe and mostly beneficial” and trying to incorporate the evidence “make them less robustly safe in various different, potentially severe ways.” There are reasonable ways to take these fact patterns and estimate the likelihood of catastrophe at much, much less than 1%, but that was not what it looked like to me.
But the tone of your comment was that I wasn’t just saying something incorrect, but that what I was doing was dangerous and counterproductive.
So was the problem was that I mentioned the fact pattern at all?
There’s an understandable story here, which says that the risk of AI takeover from Astra+ is
But I don’t think the conclusion follows. I obviously can’t know for certain if drawing attention to this kind of deployment makes the future go better, but generally my view on these things is that there’s two headline countervailing effects here:
Stopping a dangerous deployment creates procedure and precedent for stopping future dangerous deployments.
Stopping a deployment prematurely risks burning goodwill to stop more dangerous deployments in the future.
I think the first consideration matters substantially more than the second.
The fact that few people were talking about this aspect of the Navier-Stokes effort was a big reason I was more reluctant to talk about it. But I thought something was dangerous and urgent through a reasonable set of inferences and then I communicated my best understanding, in hopes that other people would be able to further look into what they found worth investigating.
In general, I think fully consequentialist justifications are intractable and that it is good to be able to point to concrete things and say “I think this could be causing harm on the world and we have a plausible theory of how to make it not happen”, even if you can point to reasons for why the second-order effects of stopping this harm might be bad.
So I think your response was unnecessarily hostile, though I hope we find friendly and collaborative ground for future conversations. Particularly I feel like it’s currently more important to evaluate the object-level claims of:
Whether the deployment should be stopped.
Whether both Millenium Prize traces should be audited, and the mechanics of doing so.
Developing a better understanding for whether this kind of deployment of pre-aligned models is a one/two-off or a recurring event.
I apologize! I agree that my tone was unnecessarily hostile. I think it’s very important to keep it friendly here. After your clarification, you weren’t at all doing what I thought. I thought you were using “loss of control” to mean something like “the model goes and does unauthorized stuff for a little while before it’s shut down” or even “the model self-exfiltrates and survives on the web and causes some minor harms”.
That would be watering down the term “loss of control” in the same way that AGI and ASI have been watered down to meaninglessness. That’s what I was objecting to, strongly. I hate arguing about terminology, but the blurring of those previous terms has done substantial harm to the discourse IMO so I wanted to prevent losing another one and making it standard practice to use established terminology to mean other things.
But you’re saying you meant the traditional usage: humans lose conjtrol of the future, permanently. We lose, unless we get insanely lucky and whatever Astra++ “wants” in the long term happens to be good for us, despite not having undergone full alignment training.
If you truly believe the risk is 1-3% then we disagree on the risk. But we don’t disagree on the use of terminology, which is what I was reacting to. Having seen AGI and ASI and “alignment” watered down to near meaninglessness, I don’t think we should allow that to happen to other terms.
I don’t think your post note was at all explict about what you meant; you didn’t define “loss of control” at all.
So I hope we can call this an honest disagreement, and I apologize for the tone and not being more explicit about exactly what I was objecting to.
And we actually don’t disagree by much. Thinking about it a little more, I’m not sure that Astra’s next-gen couldn’t self-improve a harness that would give it the memory I think LLMs need to be really dangerous, and it might have enough of the executive function and metacognitive skills that seem to be the other missing piece. I’d give it something on the order of 1%, probably lower, but that’s in the same ballpark you were referring to.
On your other point, responding to me claim that it would more likely be a useful warning shot if it got free and ran amok a bit: I agree that consequentialist reasoning is hard and fraught. But this situation is hard, so I don’t think steering by heuristics is adequate. I would personally be torn by taking that ~.5% chance of immediate takeover against the large change in attitude we’d get from a serious but failed takeover attempt. But that’s a product of just how bad I think the situation is here, that I’d consider risk on the order of parts in a thousand acceptable trades for improving odds in other scenarios, and that I’d consider them nearly negligible relative to the very large overall risk as I see it. My p(doom) is stuck at 50% based entirely on model uncertainty, because after looking at the different arguments closely, I don’t think anyone has a clue how hard alignment actually is for LLM-based AGI.
No need to respond to all of that. Thanks for your friendly engagement!