Philosophy, CS and Econ student. Researcher, still trying to figure out of what, exactly.
I like board games. Go and Mahjong are my current faves.
June Jimenez
I’ve been trying to think about this comment and I don’t fully understand it, on two layers.
Was the problem the accuracy of the events I’m worried about?
When I said loss of control, I was centrally referring to a rogue deployment causing loss of control, in the way researchers usually mean the words loss of control. I did not mean “bad outcomes”, or “harms”, I meant the thing people mean when they say loss of control. I don’t want to go into depth on any specific causal story for how a rogue deployment could make itself self-sustaining, but I think we haven’t done the work to rule out Astra+ being capable of doing so, or laying the groundwork for a future model doing so, and we’ve gotten some evidence over the past few months of these capabilities increasing. The only other things I want to note here are that the model doesn’t have to succeed immediately and directly, and that hacking is not the only surface.I agree I could have / should have been more explicit about the implicit rogue deployment → loss of control step I was making.
When I said non-negligible, I agree it would have been irresponsible of me to frame it in the way I did if my internal model had been different. But I did not mean something like 10e-6, and yes it would have been fairly irresponsible if it had been something like 10e-12. I still do think even if you find it so implausible it’s roundable to zero, it was important to point out that someone should be auditing the Navier-Stokes agent traces, for example.
But I meant something like 1-3%. This is a very hard value to estimate, because we’re drawing from the reference class “Take all of these deployments which were on net mostly safe and mostly beneficial” and trying to incorporate the evidence “make them less robustly safe in various different, potentially severe ways.” There are reasonable ways to take these fact patterns and estimate the likelihood of catastrophe at much, much less than 1%, but that was not what it looked like to me.
But the tone of your comment was that I wasn’t just saying something incorrect, but that what I was doing was dangerous and counterproductive.So was the problem was that I mentioned the fact pattern at all?
There’s an understandable story here, which says that the risk of AI takeover from Astra+ is , and the risk of broader takeover is , so even if we notice this fact pattern we shouldn’t make much noise about it.But I don’t think the conclusion follows. I obviously can’t know for certain if drawing attention to this kind of deployment makes the future go better, but generally my view on these things is that there’s two headline countervailing effects here:
Stopping a dangerous deployment creates procedure and precedent for stopping future dangerous deployments.
Stopping a deployment prematurely risks burning goodwill to stop more dangerous deployments in the future.
I think the first consideration matters substantially more than the second.
The fact that few people were talking about this aspect of the Navier-Stokes effort was a big reason I was more reluctant to talk about it. But I thought something was dangerous and urgent through a reasonable set of inferences and then I communicated my best understanding, in hopes that other people would be able to further look into what they found worth investigating.
In general, I think fully consequentialist justifications are intractable and that it is good to be able to point to concrete things and say “I think this could be causing harm on the world and we have a plausible theory of how to make it not happen”, even if you can point to reasons for why the second-order effects of stopping this harm might be bad.So I think your response was unnecessarily hostile, though I hope we find friendly and collaborative ground for future conversations. Particularly I feel like it’s currently more important to evaluate the object-level claims of:
Whether the deployment should be stopped.
Whether both Millenium Prize traces should be audited, and the mechanics of doing so.
Developing a better understanding for whether this kind of deployment of pre-aligned models is a one/two-off or a recurring event.
June Jimenez’s Shortform
OpenAI’s pursuit of Navier-Stokes seems to have been a non-negligible loss of control risk. The risk could be ongoing.
From the timeline of their description of the pursuit, they had a model that had started training on August 28th, they decided to do a preliminary run on Euler forcing on September 1st, and seem to have launched the full 10,000 agent swarm on September 3rd or 4th.
Needless to say, this model could not possibly have undergone frontier-level safety alignment and safety evaluations, consistently reported to take weeks to months.
They say “We have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics.” Astra+ seems to be in the general trend of “model capabilities emerge downstream from general scaling efforts + RLVR”, so it’s a fairly reasonable inference that Astra+ is also more capable at cyber, bio and rogue deployments.
So it seems plausible that they took a model which was:Substantially more capable than Astra
Had undergone much less prosaic alignment training than Astra
Had weaker safeguards than production-deployment Astra
Was directed to solve a problem a step-up in difficulty from those that had been previously solved and that was plausibly outside of its level of capabilities; the kind of environment that has seemed to bring out misaligned persistence recently.
And let it output 300 billion tokens in pursuit of said goal.
A couple of things:Given the gigantic extent of the research effort, shouldn’t someone, or ideally a large team of people, be auditing the agent traces to make sure that Astra+ didn’t hack a bunch of stuff en route to solving Navier-Stokes? Has someone done so?
The NYT reports that OpenAI are making substantial progress on a second Millenium Prize problem, presumably with this or a somewhat more capable version of this model. Given that it still hasn’t undergone safety training, is this wise?
More generally, I will note that models can be misaligned during training. One hopes not. How standard is it to deploy models in very early stages of training?
And should this model be trained in the first place? Astra’s training was infamously bumpy enough that Astra-class model training had to be slowed down substantially. Have you, only one to two months later, patched all of the control failure modes that would allow a more powerful system to sustain a rogue deployment?
9⁄13 Update:
Via Peter Wildeford, the “Path to Astra” blog post says that frontier RL training run was re-started on August 28th. This makes more sense in terms of timelines, but still seems to indicate deployment of a model that has not undergone alignment RL and testing. I agree with some comments about being unsure whether applying prosaic alignment practices on a frontier model are net beneficial, but a model that has undergone little safety evaluation is at least more legibly and self-evidently misaligned-by-default.
Something I’m considering. Is the Coxon moment leading to one-off interest in AI safety concerns, requiring urgent action in a narrow window, or is it part of the broader trend of increasing salience of AI in the national and global conversation?
I’m currently leaning it’s the latter (it seems non-obvious in any case), but I’m bringing this up because I feel like this matters.
Heavy on opinion, but:
I argue be urgent, but not reckless. I’m not at all arguing the speed premium for projects is 0. This may be the largest leverage moment safety researchers have had yet, and for many reasons it’s important to rapidly both continue to do good research, and to use the voice we do have to inform policymakers accurately and truthfully, advocating for good policies when advocacy is appropriate.
But there’s a big difference between “highest leverage yet” and “highest leverage ever.” Relationships between policymakers and advocates are more effective the longer term they are, and I think we should not treat this like the last moment we’ll have the public’s ears. All the effort in the world will do no good (and indeed, could do much harm) if we advocate for nonsense. Try to act, if possible, as communities; even just a second person engaging with your work, even if they’re not a formal researcher, I’ve found very helpful. Having even a couple of people critically and effortfully engaging with ideas will help defuse some of your blindspots before you present them to policymakers and not after.
And be careful when burning bridges. Burning bridges is the correct thing to do sometimes, but do it when it must be done; when it’s likely to be robustly beneficial, not just out of urgency. The sense that something needs to be done, by itself, doesn’t seem like a reliable guide to good action.