I started drafting a similar point about warning shot prevention being net-negative, which I’ll post here.
In a world where OpenAI had safer internal AI deployments, they might have ensured that their ExploitGym evals included better sandboxing, monitoring, etc. This could have prevented the HF incident from happening.
But the HF incident was a warning shot that increased OpenAI’s willingness to pay for safety, and it made the safety problems at OpenAI much clearer to the outside world. And the harms from this incident were practically nonexistent. So preventing the HF incident seems clearly net-harmful.
It seems likely that OpenAI even redirected employees from capabilities work to safety work as a result of this incident! So the time OpenAI safety employees spent improving AI control measures may have been somewhat wasted, since this work was eventually going to be done by others anyway.
Consider a model of how much a lab invests in safety, with 3 numbers:
Ideal Safety Investment (ISI): how much safety investment the lab/society at large would want if they fully understood the risks.
Targeted Safety Investment (TSI): How much safety investment the lab currently wants.
Current Safety Investment (CSI): The lab’s actual safety investment at a given time.
Whenever the lab hires a safety employee, they are attempting to get CSI closer to TSI. But if you work on AI safety, you probably think ISI is a lot higher than TSI. Maybe if you are really good at your safety job at the lab, you can punch above your weight and increase safety even more than would otherwise be expected, but it’s unlikely that you’ll approach ISI levels of safety just by “doing a really good job.”
Every time there’s a warning shot, TSI jumps up closer to ISI. Then, the lab will try to move CSI to TSI any way they can, which may route through current safety employees or any other employee they can get their hands on.
Presumably, you mostly care about CSI at the time we have really dangerous AI, not before. So maybe it makes sense to work directly on safety at a lab if you think we won’t get any warning shots in time for the lab to raise TSI and then CSI, and you think you’re much better at direct safety implementation than the marginal person the lab would have found to raise CSI. But you can just improve safety techniques and increase TSI from the outside, to help people working on safety inside the lab without substituting for them.
Helping people working on safety inside the lab still contributes to reducing warning shots. To a lesser extent, it can also reduce TSI because the company thinks it can copy work from the outside and doesn’t need to do all the safety work itself.
(I still think some research is good, but it depends on what kind of research and what timing.)
I think this is true so long as the federal government/relevant states don’t have transparency into companies around incidents, but once they do, I think warning shots become much less useful, and therefore much more worthy of prevention, and a key update that I think is warranted based on the HF incident is that the value of warning shots is not to convince the public (which currently holds an extremely strong prior around AI not being too capable, and only updates minimally off of their limited evidence they think about), but to convince elites/decision makers in government to prioritize AI safety.
More generally, one update I’ve made from the HF incident is that most of what we need to get to existential security does not rest on the public being awake and alarmed, but rather elites in the labs and governments being awake and alarmed, and in particular good AI safety (and frankly other causes which I consider to be more worthy/better like better futures work) will largely avoid focusing on convincing the public.
Once we accept that “convincing the public” matters a lot less than people who currently work on AI governance think, it follows that we should prevent warning shots once we have government-backed transparency laws.
If I follow you, you ask for less AI safety to allow a warning shot, to get more AI safety in the end ? I see how it could work the first time, but once AI safety has been increased in the lab, you should expect less warning shots. Moreover, how can you be sure to allow a mere warning shot and not full takeover ? I think we need more honeypot setups but not less control (sandboxing etc).
I started drafting a similar point about warning shot prevention being net-negative, which I’ll post here.
In a world where OpenAI had safer internal AI deployments, they might have ensured that their ExploitGym evals included better sandboxing, monitoring, etc. This could have prevented the HF incident from happening.
But the HF incident was a warning shot that increased OpenAI’s willingness to pay for safety, and it made the safety problems at OpenAI much clearer to the outside world. And the harms from this incident were practically nonexistent. So preventing the HF incident seems clearly net-harmful.
It seems likely that OpenAI even redirected employees from capabilities work to safety work as a result of this incident! So the time OpenAI safety employees spent improving AI control measures may have been somewhat wasted, since this work was eventually going to be done by others anyway.
Consider a model of how much a lab invests in safety, with 3 numbers:
Ideal Safety Investment (ISI): how much safety investment the lab/society at large would want if they fully understood the risks.
Targeted Safety Investment (TSI): How much safety investment the lab currently wants.
Current Safety Investment (CSI): The lab’s actual safety investment at a given time.
Whenever the lab hires a safety employee, they are attempting to get CSI closer to TSI. But if you work on AI safety, you probably think ISI is a lot higher than TSI. Maybe if you are really good at your safety job at the lab, you can punch above your weight and increase safety even more than would otherwise be expected, but it’s unlikely that you’ll approach ISI levels of safety just by “doing a really good job.”
Every time there’s a warning shot, TSI jumps up closer to ISI. Then, the lab will try to move CSI to TSI any way they can, which may route through current safety employees or any other employee they can get their hands on.
Presumably, you mostly care about CSI at the time we have really dangerous AI, not before. So maybe it makes sense to work directly on safety at a lab if you think we won’t get any warning shots in time for the lab to raise TSI and then CSI, and you think you’re much better at direct safety implementation than the marginal person the lab would have found to raise CSI. But you can just improve safety techniques and increase TSI from the outside, to help people working on safety inside the lab without substituting for them.
I was also thinking about this as part of evaluating what work seems useful to do. I wrote up my model here: https://www.lesswrong.com/posts/kxHiSsNh4MH82nhXD/assessing-the-impact-of-safety-work-needs-equilibrium
(My “commercial incentives equilibrium” maps to your CSI->TSI and my “risk awareness equilibrium” to TSI->ISI I think.)
Helping people working on safety inside the lab still contributes to reducing warning shots. To a lesser extent, it can also reduce TSI because the company thinks it can copy work from the outside and doesn’t need to do all the safety work itself.
(I still think some research is good, but it depends on what kind of research and what timing.)
I think this is true so long as the federal government/relevant states don’t have transparency into companies around incidents, but once they do, I think warning shots become much less useful, and therefore much more worthy of prevention, and a key update that I think is warranted based on the HF incident is that the value of warning shots is not to convince the public (which currently holds an extremely strong prior around AI not being too capable, and only updates minimally off of their limited evidence they think about), but to convince elites/decision makers in government to prioritize AI safety.
More generally, one update I’ve made from the HF incident is that most of what we need to get to existential security does not rest on the public being awake and alarmed, but rather elites in the labs and governments being awake and alarmed, and in particular good AI safety (and frankly other causes which I consider to be more worthy/better like better futures work) will largely avoid focusing on convincing the public.
Once we accept that “convincing the public” matters a lot less than people who currently work on AI governance think, it follows that we should prevent warning shots once we have government-backed transparency laws.
If I follow you, you ask for less AI safety to allow a warning shot, to get more AI safety in the end ? I see how it could work the first time, but once AI safety has been increased in the lab, you should expect less warning shots. Moreover, how can you be sure to allow a mere warning shot and not full takeover ? I think we need more honeypot setups but not less control (sandboxing etc).