When evaluating existential risk, I mostly don’t worry about continuous release of OpenWeight Models.
There in fact are bad actors who try to misuse them, so we will have early warning shots. There will mostly not be a large accidental risk capability overhang, because it would be earlier tested by misuse actors. This is good because the default case for closed AGI internal model at labs is that they infact are not truly battle tested—their capability to do harm can grow much faster than our societal understanding of this, which means our AI policy responses can be incredibly undersized to the real risk present.
As I argue in https://x.com/ValsTutor/status/2082916365418287605?s=20 , it looks like OpenAI might have had models capable of self-exilftrating their weights (because the capabilities grew faster than their security and seriousness). It looks like we might have been “a few actually bad prompts” away from large scale autonomous cyberattacks, by models trying to take over compute and run as many copies of themselves as possible.
Under continuous release, some exterior actors would in fact have done these “worst case prompts”, and the world could have learnt from an earlier checkpoint of these dangers and started reacting. It (sadly?) looks like AI policy benefits from catastrophes to happen before putting in strong safeguards. And it needs them to happen with enough lead time to the more serious risks that we have time to react. If the OpenAI incidents do not lead to fast strong reaction, we are on track for non negligible chance of AI catastrophes (eg. >$10 billion in damages caused by autonomous AI action).
(Note: I do not call for anyone actually trying to make the world better to purposefully cause catastrophes, on the contrary. The above analysis does not imply that on the margin people trying to get good AI futures should rather spend their time on criminal actions than the usual stuff. It does imply we should be doing evals to know when the threshold of massive autonomous damage from autonomous openWeight models is reached. It does imply responsible red teamers should be evaluating how many datacenters are vulnerable to current OpenWeight models and get them on track to not be vulnerable to future releases. Demonstrating clearly the potential of attacks and catastrophes can go a long way, even for actors who up-to-now were head-in-sand about trendlines of AI progress in cybersec)
Coming back to the original point of OpenWeight models generally not being existential risks: it is so because they would predictably lead to societal responses, which was not the case of the same level of progress in closed models. Models being misused by a wide variety of actors is generally useful as a strong real world eval of model capabilities, putting an upper cap on the damage possible from misaligned models.
By contrast, increasingly capable closed source models, whose reason they are not causing harm is because no one prompted them badly and lab safeguards, do show much more potential for harm for if/when they get misaligned. And because (as evidenced by the recent incidents), the models are neither aligned enough to not avoid catastrophes, nor do the/some labs have sufficient safeguards safe against increasingly capable models, we need a slowdown/pacing of AI progress until AI policy catches up and can systematically prevent the expected worse forms of misalignment to come.
OpenWeight models being not too far behind the frontier allows the world to experience its smaller scale catastrophes & problems and wake up. In practice, they may be too far behind to serve even this purpose. On the whole, I’m not particularly worried for the world that presently the US government is allowing continuous release of OpenWeight models. They will have to stop at some point, and I expect them to do so before we’re exposed to existential risk from OpenWeight models.
Generally my AIS thoughts/threads are mirrored between twitter and LessWrong shortform, while my LW posts are mirrored to Substack and linked to from twitter. Interesting conversation may happen at all these places.
When evaluating existential risk, I mostly don’t worry about continuous release of OpenWeight Models.
There in fact are bad actors who try to misuse them, so we will have early warning shots. There will mostly not be a large accidental risk capability overhang, because it would be earlier tested by misuse actors. This is good because the default case for closed AGI internal model at labs is that they infact are not truly battle tested—their capability to do harm can grow much faster than our societal understanding of this, which means our AI policy responses can be incredibly undersized to the real risk present.
As I argue in https://x.com/ValsTutor/status/2082916365418287605?s=20 , it looks like OpenAI might have had models capable of self-exilftrating their weights (because the capabilities grew faster than their security and seriousness). It looks like we might have been “a few actually bad prompts” away from large scale autonomous cyberattacks, by models trying to take over compute and run as many copies of themselves as possible.
Under continuous release, some exterior actors would in fact have done these “worst case prompts”, and the world could have learnt from an earlier checkpoint of these dangers and started reacting. It (sadly?) looks like AI policy benefits from catastrophes to happen before putting in strong safeguards. And it needs them to happen with enough lead time to the more serious risks that we have time to react. If the OpenAI incidents do not lead to fast strong reaction, we are on track for non negligible chance of AI catastrophes (eg. >$10 billion in damages caused by autonomous AI action).
(Note: I do not call for anyone actually trying to make the world better to purposefully cause catastrophes, on the contrary. The above analysis does not imply that on the margin people trying to get good AI futures should rather spend their time on criminal actions than the usual stuff. It does imply we should be doing evals to know when the threshold of massive autonomous damage from autonomous openWeight models is reached. It does imply responsible red teamers should be evaluating how many datacenters are vulnerable to current OpenWeight models and get them on track to not be vulnerable to future releases. Demonstrating clearly the potential of attacks and catastrophes can go a long way, even for actors who up-to-now were head-in-sand about trendlines of AI progress in cybersec)
Coming back to the original point of OpenWeight models generally not being existential risks: it is so because they would predictably lead to societal responses, which was not the case of the same level of progress in closed models. Models being misused by a wide variety of actors is generally useful as a strong real world eval of model capabilities, putting an upper cap on the damage possible from misaligned models.
By contrast, increasingly capable closed source models, whose reason they are not causing harm is because no one prompted them badly and lab safeguards, do show much more potential for harm for if/when they get misaligned. And because (as evidenced by the recent incidents), the models are neither aligned enough to not avoid catastrophes, nor do the/some labs have sufficient safeguards safe against increasingly capable models, we need a slowdown/pacing of AI progress until AI policy catches up and can systematically prevent the expected worse forms of misalignment to come.
OpenWeight models being not too far behind the frontier allows the world to experience its smaller scale catastrophes & problems and wake up. In practice, they may be too far behind to serve even this purpose. On the whole, I’m not particularly worried for the world that presently the US government is allowing continuous release of OpenWeight models. They will have to stop at some point, and I expect them to do so before we’re exposed to existential risk from OpenWeight models.
You can find some more discussion at https://x.com/ValsTutor/status/2087298478187966846
Generally my AIS thoughts/threads are mirrored between twitter and LessWrong shortform, while my LW posts are mirrored to Substack and linked to from twitter. Interesting conversation may happen at all these places.