The recent security incidents at OpenAI, Anthropic, and UKAISI are (a) moderate evidence that AI safety is bottlenecked by implementation of known techniques, and (b) weak evidence that AI safety is tractable*. The main reason is they would mostly have been caught by monitoring and other safeguards (OAI and Ant deliberately disabled safeguards). Some caveats:
Superintelligence may require new alignment techniques, but we also won’t have enough capacity to implement them by default.
* By “tractable” I mean that the elasticity of p(doom) to % of resources spent on safety vs capability is high, something like 10-50% relative p(doom) reduction per doubling in safety resources rather than 1%.
More speculatively, maybe futures where we narrowly avoid doom involve schemers running around at AI companies 24⁄7, but always outresourced by aligned AIs playing defense.
Watching the Black Hat video I’m struck by how this whole thing could have been prevented by not half-assing everything (impossible RL training, giving the package manager internet access, wildly inadequate monitoring, noticing the problem and then continuing to give models trained on it access to the same package manager).
More speculatively, maybe futures where we narrowly avoid doom involve schemers running around at AI companies 24⁄7, but always outresourced by aligned AIs playing defense.
Do you have a story by which there are schemers running around 24⁄7 and also the AIs playing defense are aligned?
Aligned AIs take 2x the inference compute to run as unaligned AIs, but companies don’t know the exact threshold or which mitigations they need, so they constantly experiment on checkpoints. They ultimately need to pick 10 of 50 different safety techniques, each of which halves misalignment rate from a baseline of 100% and costs 10% of baseline inference compute. It’s more sample efficient to test the AIs with only one mitigation applied at a time, and it gives more signal to test them in realistic internal deployments. Safety teams are very nervous about the defensive AIs being so load-bearing, but empirically it works well and this lets them get more safety research done.
Similar to today, companies have inadequate operational practices and don’t reliably apply even cheap safeguards. 90% of research is automated so spending human labor to improve them would be an extreme bottleneck. Instead, they give aligned AIs latitude to stop unaligned AIs however they see fit, which leads to a disorganized patchwork of shallow alignment, control, etc. Companies are large enough that on average, at least one schemer at a time makes it through the first 5 layers of Swiss cheese, and there are a few more layers before doom happens. The defensive AIs are probably not CEV aligned, they’re just non-power-seeking and aligned enough to be used as tools.
I agree if you mean this in the sense of, someone infecting you with HIV suddenly makes it highly tractable extend your own lifespan.
Like my updates from these incidents are:
Weak update on
(downwards) How well current approaches are working for sub-superhuman capability levels
(upwards) How much effort we need to put in to get them to work
Moderate negative update on a mix of
Overall lab competence
How seriously big labs take misalignment risks
How rushed and hectic it is inside big labs
Like, I’m learning a problem I thought was relatively easy(*), is actually harder. And the people I thought were on top of the problem are actually not at all.
And this does mean there is more room for improvement. But its still straightforwardly a negative update on the overall situation (modulo things like this scaring people / waking people up).
*easy in the sense, I’d strongly suspect smart hard-working people be able to solve it
The recent security incidents at OpenAI, Anthropic, and UKAISI are (a) moderate evidence that AI safety is bottlenecked by implementation of known techniques, and (b) weak evidence that AI safety is tractable*. The main reason is they would mostly have been caught by monitoring and other safeguards (OAI and Ant deliberately disabled safeguards). Some caveats:
Superintelligence may require new alignment techniques, but we also won’t have enough capacity to implement them by default.
* By “tractable” I mean that the elasticity of p(doom) to % of resources spent on safety vs capability is high, something like 10-50% relative p(doom) reduction per doubling in safety resources rather than 1%.
More speculatively, maybe futures where we narrowly avoid doom involve schemers running around at AI companies 24⁄7, but always outresourced by aligned AIs playing defense.
Watching the Black Hat video I’m struck by how this whole thing could have been prevented by not half-assing everything (impossible RL training, giving the package manager internet access, wildly inadequate monitoring, noticing the problem and then continuing to give models trained on it access to the same package manager).
https://youtu.be/87DyyMV0kCY?is=fJuUyk597v8ibE1v
Do you have a story by which there are schemers running around 24⁄7 and also the AIs playing defense are aligned?
Sure, here are two possible scenarios.
Aligned AIs take 2x the inference compute to run as unaligned AIs, but companies don’t know the exact threshold or which mitigations they need, so they constantly experiment on checkpoints. They ultimately need to pick 10 of 50 different safety techniques, each of which halves misalignment rate from a baseline of 100% and costs 10% of baseline inference compute. It’s more sample efficient to test the AIs with only one mitigation applied at a time, and it gives more signal to test them in realistic internal deployments. Safety teams are very nervous about the defensive AIs being so load-bearing, but empirically it works well and this lets them get more safety research done.
Similar to today, companies have inadequate operational practices and don’t reliably apply even cheap safeguards. 90% of research is automated so spending human labor to improve them would be an extreme bottleneck. Instead, they give aligned AIs latitude to stop unaligned AIs however they see fit, which leads to a disorganized patchwork of shallow alignment, control, etc. Companies are large enough that on average, at least one schemer at a time makes it through the first 5 layers of Swiss cheese, and there are a few more layers before doom happens. The defensive AIs are probably not CEV aligned, they’re just non-power-seeking and aligned enough to be used as tools.
I agree if you mean this in the sense of, someone infecting you with HIV suddenly makes it highly tractable extend your own lifespan.
Like my updates from these incidents are:
Weak update on
(downwards) How well current approaches are working for sub-superhuman capability levels
(upwards) How much effort we need to put in to get them to work
Moderate negative update on a mix of
Overall lab competence
How seriously big labs take misalignment risks
How rushed and hectic it is inside big labs
Like, I’m learning a problem I thought was relatively easy(*), is actually harder. And the people I thought were on top of the problem are actually not at all.
And this does mean there is more room for improvement. But its still straightforwardly a negative update on the overall situation (modulo things like this scaring people / waking people up).
*easy in the sense, I’d strongly suspect smart hard-working people be able to solve it
If you replace “AI safety” with “AI control”, I think more would agree. Also +1 to the ASI caveat.