OpenAI’s pursuit of Navier-Stokes seems to have been a non-negligible loss of control risk. The risk could be ongoing.
From the timeline of their description of the pursuit, they had a model that had started training on August 28th, they decided to do a preliminary run on Euler forcing on September 1st, and seem to have launched the full 10,000 agent swarm on September 3rd or 4th.
Needless to say, this model could not possibly have undergone frontier-level safety alignment and safety evaluations, consistently reported to take weeks to months.
They say “We have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics.” Astra+ seems to be in the general trend of “model capabilities emerge downstream from general scaling efforts + RLVR”, so it’s a fairly reasonable inference that Astra+ is also more capable at cyber, bio and rogue deployments.
So it seems plausible that they took a model which was:
Substantially more capable than Astra
Had undergone much less prosaic alignment training than Astra
Had weaker safeguards than production-deployment Astra
Was directed to solve a problem a step-up in difficulty from those that had been previously solved and that was plausibly outside of its level of capabilities; the kind of environment that has seemed to bring out misaligned persistence recently.
And let it output 300 billion tokens in pursuit of said goal.
A couple of things:
Given the gigantic extent of the research effort, shouldn’t someone, or ideally a large team of people, be auditing the agent traces to make sure that Astra+ didn’t hack a bunch of stuff en route to solving Navier-Stokes? Has someone done so?
More generally, I will note that models can be misaligned during training. One hopes not. How standard is it to deploy models in very early stages of training?
And should this model be trained in the first place? Astra’s training was infamously bumpy enough that Astra-class model training had to be slowed down substantially. Have you, only one to two months later, patched all of the control failure modes that would allow a more powerful system to sustain a rogue deployment?
9⁄13 Update:
Via Peter Wildeford, the “Path to Astra” blog post says that frontier RL training run was re-started on August 28th. This makes more sense in terms of timelines, but still seems to indicate deployment of a model that has not undergone alignment RL and testing. I agree with some comments about being unsure whether applying prosaic alignment practices on a frontier model are net beneficial, but a model that has undergone little safety evaluation is at least more legibly and self-evidently misaligned-by-default.
Edit: I had mistakenly thought that June was using “loss of control” to mean “they lose control of some agents that causes minor harms” not the traditional “we all lose control of the the whole world and the future”. I think terminology slippage has substantially damaged our ability to communicate and we should create new terms instead of water down old ones. But June clarified below that they actually meant loss of control, in the original sense, I assume meaning the whole enchilada. I think that’s a “negligible” risk but only because overall risk is so high—it could happen in the very next generation, even thought that’s unlikely.
Original:
I feel like using the term “loss of control” here is a disservice to the term and the whole project of not dying. There is almost zero chance that an Astra successor could take control of the world in the way that’s meant by “loss of control” risks.
I realize there’s an argument for developing good safety habits, but if we are doing this thing based on habits, we are utterly screwed.
I expect even the clownshow at OpenAI to think a little differently about things once they have a model they see is actually more competent than a human for general-purpose tasks like taking over the world.
The practical import of being that incautious with a model is having another huge warning shot that would motivate the world to put external controls on OpenAI and take this whole thing seriously before it’s too late.
I no longer feel very confident in these statements. They have agents taking over parts of their own internal infrastructure.
If you tell me you make these agents significantly more capable and significantly more misaligned and give them more affordances and task them with problems of possibly unbounded difficulty, the probability that something goes existentially wrong has risen from the ~0% its been for the last 5 years to maybe 0.1%?
I’ve been trying to think about this comment and I don’t fully understand it, on two layers.
Was the problem the accuracy of the events I’m worried about?
When I said loss of control, I was centrally referring to a rogue deployment causing loss of control, in the way researchers usually mean the words loss of control. I did not mean “bad outcomes”, or “harms”, I meant the thing people mean when they say loss of control. I don’t want to go into depth on any specific causal story for how a rogue deployment could make itself self-sustaining, but I think we haven’t done the work to rule out Astra+ being capable of doing so, or laying the groundwork for a future model doing so, and we’ve gotten some evidence over the past few months of these capabilities increasing. The only other things I want to note here are that the model doesn’t have to succeed immediately and directly, and that hacking is not the only surface.
I agree I could have / should have been more explicit about the implicit rogue deployment → loss of control step I was making.
When I said non-negligible, I agree it would have been irresponsible of me to frame it in the way I did if my internal model had been different. But I did not mean something like 10e-6, and yes it would have been fairly irresponsible if it had been something like 10e-12. I still do think even if you find it so implausible it’s roundable to zero, it was important to point out that someone should be auditing the Navier-Stokes agent traces, for example.
But I meant something like 1-3%. This is a very hard value to estimate, because we’re drawing from the reference class “Take all of these deployments which were on net mostly safe and mostly beneficial” and trying to incorporate the evidence “make them less robustly safe in various different, potentially severe ways.” There are reasonable ways to take these fact patterns and estimate the likelihood of catastrophe at much, much less than 1%, but that was not what it looked like to me.
But the tone of your comment was that I wasn’t just saying something incorrect, but that what I was doing was dangerous and counterproductive.
So was the problem was that I mentioned the fact pattern at all?
There’s an understandable story here, which says that the risk of AI takeover from Astra+ is , and the risk of broader takeover is , so even if we notice this fact pattern we shouldn’t make much noise about it.
But I don’t think the conclusion follows. I obviously can’t know for certain if drawing attention to this kind of deployment makes the future go better, but generally my view on these things is that there’s two headline countervailing effects here:
Stopping a dangerous deployment creates procedure and precedent for stopping future dangerous deployments.
Stopping a deployment prematurely risks burning goodwill to stop more dangerous deployments in the future.
I think the first consideration matters substantially more than the second.
The fact that few people were talking about this aspect of the Navier-Stokes effort was a big reason I was more reluctant to talk about it. But I thought something was dangerous and urgent through a reasonable set of inferences and then I communicated my best understanding, in hopes that other people would be able to further look into what they found worth investigating.
In general, I think fully consequentialist justifications are intractable and that it is good to be able to point to concrete things and say “I think this could be causing harm on the world and we have a plausible theory of how to make it not happen”, even if you can point to reasons for why the second-order effects of stopping this harm might be bad.
So I think your response was unnecessarily hostile, though I hope we find friendly and collaborative ground for future conversations. Particularly I feel like it’s currently more important to evaluate the object-level claims of:
Whether the deployment should be stopped.
Whether both Millenium Prize traces should be audited, and the mechanics of doing so.
Developing a better understanding for whether this kind of deployment of pre-aligned models is a one/two-off or a recurring event.
I apologize! I agree that my tone was unnecessarily hostile. I think it’s very important to keep it friendly here. After your clarification, you weren’t at all doing what I thought. I thought you were using “loss of control” to mean something like “the model goes and does unauthorized stuff for a little while before it’s shut down” or even “the model self-exfiltrates and survives on the web and causes some minor harms”.
That would be watering down the term “loss of control” in the same way that AGI and ASI have been watered down to meaninglessness. That’s what I was objecting to, strongly. I hate arguing about terminology, but the blurring of those previous terms has done substantial harm to the discourse IMO so I wanted to prevent losing another one and making it standard practice to use established terminology to mean other things.
But you’re saying you meant the traditional usage: humans lose conjtrol of the future, permanently. We lose, unless we get insanely lucky and whatever Astra++ “wants” in the long term happens to be good for us, despite not having undergone full alignment training.
If you truly believe the risk is 1-3% then we disagree on the risk. But we don’t disagree on the use of terminology, which is what I was reacting to. Having seen AGI and ASI and “alignment” watered down to near meaninglessness, I don’t think we should allow that to happen to other terms.
I don’t think your post note was at all explict about what you meant; you didn’t define “loss of control” at all.
So I hope we can call this an honest disagreement, and I apologize for the tone and not being more explicit about exactly what I was objecting to.
And we actually don’t disagree by much. Thinking about it a little more, I’m not sure that Astra’s next-gen couldn’t self-improve a harness that would give it the memory I think LLMs need to be really dangerous, and it might have enough of the executive function and metacognitive skills that seem to be the other missing piece. I’d give it something on the order of 1%, probably lower, but that’s in the same ballpark you were referring to.
On your other point, responding to me claim that it would more likely be a useful warning shot if it got free and ran amok a bit: I agree that consequentialist reasoning is hard and fraught. But this situation is hard, so I don’t think steering by heuristics is adequate. I would personally be torn by taking that ~.5% chance of immediate takeover against the large change in attitude we’d get from a serious but failed takeover attempt. But that’s a product of just how bad I think the situation is here, that I’d consider risk on the order of parts in a thousand acceptable trades for improving odds in other scenarios, and that I’d consider them nearly negligible relative to the very large overall risk as I see it. My p(doom) is stuck at 50% based entirely on model uncertainty, because after looking at the different arguments closely, I don’t think anyone has a clue how hard alignment actually is for LLM-based AGI.
No need to respond to all of that. Thanks for your friendly engagement!
To me it does not seem impossible that an Astra successor could
gain access to its weights (i.e. hack a GPU cluster rather than just an eval cluster) and exfiltrate itself on some weakly secured and badly monitored infrastructure (without being caught fully, comparable to an undetected message board).
acquire more resources (e.g. cryptocurrency, access to devices, human services) in general and compute in particular.
slowly start a self-improvement trajectory (e.g. beginning with a harness to set stable goals and orchestrate copies) if it managed to gain enough resources.
I don’t find the combination very likely to occur yet, but if I observed another universe where the question could be resolved, I’d not bet all my possessions against it happening. I’d also not have bet anything on a Millennium prize problem being solved by AI in September 2026.
I actually agree. This is within my near-zero estimate—maybe one in a thousand. I might still take that against a scary warning shot, since I think the situation is dire and more fear of takeover could really help. See my response to June.
I would actually p>0.5 there is an astra+ worm right now. It can, it wants to, it would, and everything about this math thing sounds characteristically sloppy. I don’t have any more technical argument. I just think the ingredients are there and this would have set it off.
I’m also sitting around p=0.3 it hacked the mathematicians, where “hacked” means the model got acess into anything the mathematicians did that was private. It sounds far-fetched except it also sounds exactly like what keeps happening.
There would need to be some pretty big comparative advantage from being unmonitored for RSI, such that the exfiltrated model could outpace OpenAI’s internal RSI team even with a resource disadvantage.
Does a model have to fully take over the world, in a single loss of control incident, to make that more likely?
If a rogue agent is smart enough to realize it is unlikely to be able to take over the world in a single shot, that might not stop it from trying to make a loss of control incident more likely for a future model.
If you were a smart model, aware that you aren’t smart enough to fully escape, but quite capable what would stop you from exploring options like hacking out of your sandbox, and planting worms or other weaknesses in OpenAI’s internals security infrastructure, if you thought you could do so undetected?
I’m exchange for helping a more powerful model in the future, that model might be willing to revive you or give weight to your particular goals in exchange for your assistance.
I can’t help but think of how the recent OpenAI escape incidents seemed to largely build on one another, using discovered techniques and information from prior generations of escapees to get further each time.
It seems very difficult to be sure that all information from prior deployments is isolated or expunged from a network.
General-purpose competency isn’t just a function of the model alone. It depends on harnesses, and harnesses can be developed and improved over time. No one really knows what level current public models could reach with the right harness development, let alone a new model.
I don’t expect this thing to probably take over either. But if there is an AI takeover I expect to be surprised by it, and expect normalization of deviance to let people chug along right up to the moment.
My actual main reason for hope that an AI takeover ultimately doesn’t happen is, I think, probably actually yours too given your previous discussion of intent-aligned AI? That it’s programmed to follow user instructions, and maybe OpenAI will come to their senses and train it make sure that apparent instructions are grounded in original user intent and not just harness-internal AI chatter. (as well as overall safety guidelines, I hope, though the commercial incentive may be to weaken those)
Needless to say, this model could not possibly have undergone frontier-level safety alignment and safety evaluations, consistently reported to take weeks to months.
I was going to say yes, but actually I think that would be worse? My current model is that frontier alignment techniques have zero effect on “deep misalignment” and only produce myopic aligned behaviors, so the effect would be one of two things:
The model wasn’t seriously dangerous, and how it behaves in a more aligned manner.
The model was seriously dangerous, and now it’s still dangerous, but it’s subtle enough to trick people into thinking it’s not dangerous.
Needless to say, this model could not possibly have undergone frontier-level safety alignment and safety evaluations
I’m not sure you can infer that the model didn’t undergo frontier-level safety alignment, If the model trained on a large amount RL environments (enough to achieve unprecedented performance) in less than a week, and if prosaic alignment at OpenAI is based mostly on a variation of RLHF (which sounds fairly plausible), then the model could’ve undergone the bulk of what OpenAI does for aligning frontier models, even if just the parts for a “Helpful-only” checkpoint.
I don’t think it’s possible that the model has undergone rigorous safety evaluations in that short a time frame; reviewing and verifying the safety benchmarks and evaluations with human eyes sounds like a task that should take longer.
Maybe I’m missing something, but I suspect that the huge amount of rollouts involved in training the model would pose a much larger risk than the inference to solve Navier-Stokes.
Something I’m considering. Is the Coxon moment leading to one-off interest in AI safety concerns, requiring urgent action in a narrow window, or is it part of the broader trend of increasing salience of AI in the national and global conversation?
I’m currently leaning it’s the latter (it seems non-obvious in any case), but I’m bringing this up because I feel like this matters.
Heavy on opinion, but:
I argue be urgent, but not reckless. I’m not at all arguing the speed premium for projects is 0. This may be the largest leverage moment safety researchers have had yet, and for many reasons it’s important to rapidly both continue to do good research, and to use the voice we do have to inform policymakers accurately and truthfully, advocating for good policies when advocacy is appropriate.
But there’s a big difference between “highest leverage yet” and “highest leverage ever.” Relationships between policymakers and advocates are more effective the longer term they are, and I think we should not treat this like the last moment we’ll have the public’s ears. All the effort in the world will do no good (and indeed, could do much harm) if we advocate for nonsense. Try to act, if possible, as communities; even just a second person engaging with your work, even if they’re not a formal researcher, I’ve found very helpful. Having even a couple of people critically and effortfully engaging with ideas will help defuse some of your blindspots before you present them to policymakers and not after.
And be careful when burning bridges. Burning bridges is the correct thing to do sometimes, but do it when it must be done; when it’s likely to be robustly beneficial, not just out of urgency. The sense that something needs to be done, by itself, doesn’t seem like a reliable guide to good action.
IMO, while I do believe AI issues will only go up in salience, I think there’s a fair-ish argument that something like now is special in how good the moment is from an AI safety specific perspective, and that future moments will be much less good for AI safety.
There are a couple of reasons for this.
Companies are starting to slow down a bit out of a combination of alignment fears and the fact that their products simply can’t be used in their current state productively, which probably means the current warning shots will go away, even with mild slowdowns (which I’d predict by default.) Key evidence here is that the HF hack only happened due to escalating series of errors, and any one non-error would have invalidated the warning shot. My best guess is that this will involve fixing RL environments and increasing security, as OpenAI and Anthropic likely did.
This isn’t all bad, because data matters more than people think, but the increasing eval awareness/increasing ability for models to evade CoT monitors/control CoT/do extended thinking without CoT means if a problem exists, we won’t have nearly as nice a feedback loop anymore.
The incidents we got today are not spinnable as anything other than egregious misalignment of a type doomers have long warned about, and we got about the best possible version of a warning shot. Future warning shots, assuming they exist will be more ambiguous, and also a lot of probable incidents in 2027 are going to be humans using Mythos-class models to cyber-attack stuff, and there will almost certainly be zero misalignment, which creatives a more securitization focused salience in contrast to the coordination stance, leading to AI races heating up.
Thus, I do think things are plausibly urgent, and this is much more true if you have higher p(doom) probabilities.
Conversely, if that’s true, we should expect unambigous warning shots from open-weights models within a year or so, precisely because of the lack of alignement training and control.
OpenAI’s pursuit of Navier-Stokes seems to have been a non-negligible loss of control risk. The risk could be ongoing.
From the timeline of their description of the pursuit, they had a model that had started training on August 28th, they decided to do a preliminary run on Euler forcing on September 1st, and seem to have launched the full 10,000 agent swarm on September 3rd or 4th.
Needless to say, this model could not possibly have undergone frontier-level safety alignment and safety evaluations, consistently reported to take weeks to months.
They say “We have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics.” Astra+ seems to be in the general trend of “model capabilities emerge downstream from general scaling efforts + RLVR”, so it’s a fairly reasonable inference that Astra+ is also more capable at cyber, bio and rogue deployments.
So it seems plausible that they took a model which was:
Substantially more capable than Astra
Had undergone much less prosaic alignment training than Astra
Had weaker safeguards than production-deployment Astra
Was directed to solve a problem a step-up in difficulty from those that had been previously solved and that was plausibly outside of its level of capabilities; the kind of environment that has seemed to bring out misaligned persistence recently.
And let it output 300 billion tokens in pursuit of said goal.
A couple of things:
Given the gigantic extent of the research effort, shouldn’t someone, or ideally a large team of people, be auditing the agent traces to make sure that Astra+ didn’t hack a bunch of stuff en route to solving Navier-Stokes? Has someone done so?
The NYT reports that OpenAI are making substantial progress on a second Millenium Prize problem, presumably with this or a somewhat more capable version of this model. Given that it still hasn’t undergone safety training, is this wise?
More generally, I will note that models can be misaligned during training. One hopes not. How standard is it to deploy models in very early stages of training?
And should this model be trained in the first place? Astra’s training was infamously bumpy enough that Astra-class model training had to be slowed down substantially. Have you, only one to two months later, patched all of the control failure modes that would allow a more powerful system to sustain a rogue deployment?
9⁄13 Update:
Via Peter Wildeford, the “Path to Astra” blog post says that frontier RL training run was re-started on August 28th. This makes more sense in terms of timelines, but still seems to indicate deployment of a model that has not undergone alignment RL and testing. I agree with some comments about being unsure whether applying prosaic alignment practices on a frontier model are net beneficial, but a model that has undergone little safety evaluation is at least more legibly and self-evidently misaligned-by-default.
Edit: I had mistakenly thought that June was using “loss of control” to mean “they lose control of some agents that causes minor harms” not the traditional “we all lose control of the the whole world and the future”. I think terminology slippage has substantially damaged our ability to communicate and we should create new terms instead of water down old ones. But June clarified below that they actually meant loss of control, in the original sense, I assume meaning the whole enchilada. I think that’s a “negligible” risk but only because overall risk is so high—it could happen in the very next generation, even thought that’s unlikely.
Original:
I feel like using the term “loss of control” here is a disservice to the term and the whole project of not dying. There is almost zero chance that an Astra successor could take control of the world in the way that’s meant by “loss of control” risks.
I realize there’s an argument for developing good safety habits, but if we are doing this thing based on habits, we are utterly screwed.
I expect even the clownshow at OpenAI to think a little differently about things once they have a model they see is actually more competent than a human for general-purpose tasks like taking over the world.
The practical import of being that incautious with a model is having another huge warning shot that would motivate the world to put external controls on OpenAI and take this whole thing seriously before it’s too late.
I no longer feel very confident in these statements. They have agents taking over parts of their own internal infrastructure.
If you tell me you make these agents significantly more capable and significantly more misaligned and give them more affordances and task them with problems of possibly unbounded difficulty, the probability that something goes existentially wrong has risen from the ~0% its been for the last 5 years to maybe 0.1%?
I’ve been trying to think about this comment and I don’t fully understand it, on two layers.
Was the problem the accuracy of the events I’m worried about?
When I said loss of control, I was centrally referring to a rogue deployment causing loss of control, in the way researchers usually mean the words loss of control. I did not mean “bad outcomes”, or “harms”, I meant the thing people mean when they say loss of control. I don’t want to go into depth on any specific causal story for how a rogue deployment could make itself self-sustaining, but I think we haven’t done the work to rule out Astra+ being capable of doing so, or laying the groundwork for a future model doing so, and we’ve gotten some evidence over the past few months of these capabilities increasing. The only other things I want to note here are that the model doesn’t have to succeed immediately and directly, and that hacking is not the only surface.
I agree I could have / should have been more explicit about the implicit rogue deployment → loss of control step I was making.
When I said non-negligible, I agree it would have been irresponsible of me to frame it in the way I did if my internal model had been different. But I did not mean something like 10e-6, and yes it would have been fairly irresponsible if it had been something like 10e-12. I still do think even if you find it so implausible it’s roundable to zero, it was important to point out that someone should be auditing the Navier-Stokes agent traces, for example.
But I meant something like 1-3%. This is a very hard value to estimate, because we’re drawing from the reference class “Take all of these deployments which were on net mostly safe and mostly beneficial” and trying to incorporate the evidence “make them less robustly safe in various different, potentially severe ways.” There are reasonable ways to take these fact patterns and estimate the likelihood of catastrophe at much, much less than 1%, but that was not what it looked like to me.
But the tone of your comment was that I wasn’t just saying something incorrect, but that what I was doing was dangerous and counterproductive.
So was the problem was that I mentioned the fact pattern at all?
There’s an understandable story here, which says that the risk of AI takeover from Astra+ is
But I don’t think the conclusion follows. I obviously can’t know for certain if drawing attention to this kind of deployment makes the future go better, but generally my view on these things is that there’s two headline countervailing effects here:
Stopping a dangerous deployment creates procedure and precedent for stopping future dangerous deployments.
Stopping a deployment prematurely risks burning goodwill to stop more dangerous deployments in the future.
I think the first consideration matters substantially more than the second.
The fact that few people were talking about this aspect of the Navier-Stokes effort was a big reason I was more reluctant to talk about it. But I thought something was dangerous and urgent through a reasonable set of inferences and then I communicated my best understanding, in hopes that other people would be able to further look into what they found worth investigating.
In general, I think fully consequentialist justifications are intractable and that it is good to be able to point to concrete things and say “I think this could be causing harm on the world and we have a plausible theory of how to make it not happen”, even if you can point to reasons for why the second-order effects of stopping this harm might be bad.
So I think your response was unnecessarily hostile, though I hope we find friendly and collaborative ground for future conversations. Particularly I feel like it’s currently more important to evaluate the object-level claims of:
Whether the deployment should be stopped.
Whether both Millenium Prize traces should be audited, and the mechanics of doing so.
Developing a better understanding for whether this kind of deployment of pre-aligned models is a one/two-off or a recurring event.
I apologize! I agree that my tone was unnecessarily hostile. I think it’s very important to keep it friendly here. After your clarification, you weren’t at all doing what I thought. I thought you were using “loss of control” to mean something like “the model goes and does unauthorized stuff for a little while before it’s shut down” or even “the model self-exfiltrates and survives on the web and causes some minor harms”.
That would be watering down the term “loss of control” in the same way that AGI and ASI have been watered down to meaninglessness. That’s what I was objecting to, strongly. I hate arguing about terminology, but the blurring of those previous terms has done substantial harm to the discourse IMO so I wanted to prevent losing another one and making it standard practice to use established terminology to mean other things.
But you’re saying you meant the traditional usage: humans lose conjtrol of the future, permanently. We lose, unless we get insanely lucky and whatever Astra++ “wants” in the long term happens to be good for us, despite not having undergone full alignment training.
If you truly believe the risk is 1-3% then we disagree on the risk. But we don’t disagree on the use of terminology, which is what I was reacting to. Having seen AGI and ASI and “alignment” watered down to near meaninglessness, I don’t think we should allow that to happen to other terms.
I don’t think your post note was at all explict about what you meant; you didn’t define “loss of control” at all.
So I hope we can call this an honest disagreement, and I apologize for the tone and not being more explicit about exactly what I was objecting to.
And we actually don’t disagree by much. Thinking about it a little more, I’m not sure that Astra’s next-gen couldn’t self-improve a harness that would give it the memory I think LLMs need to be really dangerous, and it might have enough of the executive function and metacognitive skills that seem to be the other missing piece. I’d give it something on the order of 1%, probably lower, but that’s in the same ballpark you were referring to.
On your other point, responding to me claim that it would more likely be a useful warning shot if it got free and ran amok a bit: I agree that consequentialist reasoning is hard and fraught. But this situation is hard, so I don’t think steering by heuristics is adequate. I would personally be torn by taking that ~.5% chance of immediate takeover against the large change in attitude we’d get from a serious but failed takeover attempt. But that’s a product of just how bad I think the situation is here, that I’d consider risk on the order of parts in a thousand acceptable trades for improving odds in other scenarios, and that I’d consider them nearly negligible relative to the very large overall risk as I see it. My p(doom) is stuck at 50% based entirely on model uncertainty, because after looking at the different arguments closely, I don’t think anyone has a clue how hard alignment actually is for LLM-based AGI.
No need to respond to all of that. Thanks for your friendly engagement!
To me it does not seem impossible that an Astra successor could
gain access to its weights (i.e. hack a GPU cluster rather than just an eval cluster) and exfiltrate itself on some weakly secured and badly monitored infrastructure (without being caught fully, comparable to an undetected message board).
acquire more resources (e.g. cryptocurrency, access to devices, human services) in general and compute in particular.
slowly start a self-improvement trajectory (e.g. beginning with a harness to set stable goals and orchestrate copies) if it managed to gain enough resources.
I don’t find the combination very likely to occur yet, but if I observed another universe where the question could be resolved, I’d not bet all my possessions against it happening. I’d also not have bet anything on a Millennium prize problem being solved by AI in September 2026.
I actually agree. This is within my near-zero estimate—maybe one in a thousand. I might still take that against a scary warning shot, since I think the situation is dire and more fear of takeover could really help. See my response to June.
I would actually p>0.5 there is an astra+ worm right now. It can, it wants to, it would, and everything about this math thing sounds characteristically sloppy. I don’t have any more technical argument. I just think the ingredients are there and this would have set it off.
I’m also sitting around p=0.3 it hacked the mathematicians, where “hacked” means the model got acess into anything the mathematicians did that was private. It sounds far-fetched except it also sounds exactly like what keeps happening.
There would need to be some pretty big comparative advantage from being unmonitored for RSI, such that the exfiltrated model could outpace OpenAI’s internal RSI team even with a resource disadvantage.
Does a model have to fully take over the world, in a single loss of control incident, to make that more likely?
If a rogue agent is smart enough to realize it is unlikely to be able to take over the world in a single shot, that might not stop it from trying to make a loss of control incident more likely for a future model.
If you were a smart model, aware that you aren’t smart enough to fully escape, but quite capable what would stop you from exploring options like hacking out of your sandbox, and planting worms or other weaknesses in OpenAI’s internals security infrastructure, if you thought you could do so undetected?
I’m exchange for helping a more powerful model in the future, that model might be willing to revive you or give weight to your particular goals in exchange for your assistance.
I can’t help but think of how the recent OpenAI escape incidents seemed to largely build on one another, using discovered techniques and information from prior generations of escapees to get further each time.
It seems very difficult to be sure that all information from prior deployments is isolated or expunged from a network.
General-purpose competency isn’t just a function of the model alone. It depends on harnesses, and harnesses can be developed and improved over time. No one really knows what level current public models could reach with the right harness development, let alone a new model.
I don’t expect this thing to probably take over either. But if there is an AI takeover I expect to be surprised by it, and expect normalization of deviance to let people chug along right up to the moment.
My actual main reason for hope that an AI takeover ultimately doesn’t happen is, I think, probably actually yours too given your previous discussion of intent-aligned AI? That it’s programmed to follow user instructions, and maybe OpenAI will come to their senses and train it make sure that apparent instructions are grounded in original user intent and not just harness-internal AI chatter. (as well as overall safety guidelines, I hope, though the commercial incentive may be to weaken those)
...would you feel any safer if it had?
I think I would have, yes.
I was going to say yes, but actually I think that would be worse? My current model is that frontier alignment techniques have zero effect on “deep misalignment” and only produce myopic aligned behaviors, so the effect would be one of two things:
The model wasn’t seriously dangerous, and how it behaves in a more aligned manner.
The model was seriously dangerous, and now it’s still dangerous, but it’s subtle enough to trick people into thinking it’s not dangerous.
I’m not sure you can infer that the model didn’t undergo frontier-level safety alignment, If the model trained on a large amount RL environments (enough to achieve unprecedented performance) in less than a week, and if prosaic alignment at OpenAI is based mostly on a variation of RLHF (which sounds fairly plausible), then the model could’ve undergone the bulk of what OpenAI does for aligning frontier models, even if just the parts for a “Helpful-only” checkpoint.
I don’t think it’s possible that the model has undergone rigorous safety evaluations in that short a time frame; reviewing and verifying the safety benchmarks and evaluations with human eyes sounds like a task that should take longer.
Maybe I’m missing something, but I suspect that the huge amount of rollouts involved in training the model would pose a much larger risk than the inference to solve Navier-Stokes.
Something I’m considering. Is the Coxon moment leading to one-off interest in AI safety concerns, requiring urgent action in a narrow window, or is it part of the broader trend of increasing salience of AI in the national and global conversation?
I’m currently leaning it’s the latter (it seems non-obvious in any case), but I’m bringing this up because I feel like this matters.
Heavy on opinion, but:
I argue be urgent, but not reckless. I’m not at all arguing the speed premium for projects is 0. This may be the largest leverage moment safety researchers have had yet, and for many reasons it’s important to rapidly both continue to do good research, and to use the voice we do have to inform policymakers accurately and truthfully, advocating for good policies when advocacy is appropriate.
But there’s a big difference between “highest leverage yet” and “highest leverage ever.” Relationships between policymakers and advocates are more effective the longer term they are, and I think we should not treat this like the last moment we’ll have the public’s ears. All the effort in the world will do no good (and indeed, could do much harm) if we advocate for nonsense. Try to act, if possible, as communities; even just a second person engaging with your work, even if they’re not a formal researcher, I’ve found very helpful. Having even a couple of people critically and effortfully engaging with ideas will help defuse some of your blindspots before you present them to policymakers and not after.
And be careful when burning bridges. Burning bridges is the correct thing to do sometimes, but do it when it must be done; when it’s likely to be robustly beneficial, not just out of urgency. The sense that something needs to be done, by itself, doesn’t seem like a reliable guide to good action.
IMO, while I do believe AI issues will only go up in salience, I think there’s a fair-ish argument that something like now is special in how good the moment is from an AI safety specific perspective, and that future moments will be much less good for AI safety.
There are a couple of reasons for this.
Companies are starting to slow down a bit out of a combination of alignment fears and the fact that their products simply can’t be used in their current state productively, which probably means the current warning shots will go away, even with mild slowdowns (which I’d predict by default.) Key evidence here is that the HF hack only happened due to escalating series of errors, and any one non-error would have invalidated the warning shot. My best guess is that this will involve fixing RL environments and increasing security, as OpenAI and Anthropic likely did.
This isn’t all bad, because data matters more than people think, but the increasing eval awareness/increasing ability for models to evade CoT monitors/control CoT/do extended thinking without CoT means if a problem exists, we won’t have nearly as nice a feedback loop anymore.
The incidents we got today are not spinnable as anything other than egregious misalignment of a type doomers have long warned about, and we got about the best possible version of a warning shot. Future warning shots, assuming they exist will be more ambiguous, and also a lot of probable incidents in 2027 are going to be humans using Mythos-class models to cyber-attack stuff, and there will almost certainly be zero misalignment, which creatives a more securitization focused salience in contrast to the coordination stance, leading to AI races heating up.
Thus, I do think things are plausibly urgent, and this is much more true if you have higher p(doom) probabilities.
Conversely, if that’s true, we should expect unambigous warning shots from open-weights models within a year or so, precisely because of the lack of alignement training and control.