I think the main motivation of why warning shots seem good isn’t that labs take safety more seriously but that the government does. Like if a model had exfiltrated itself and hacked more datacenters and drained cryptocurrencies and we had to scramble to shut it off, the USG might start to consider things like Plan A.
sure, i think we actually agree.
I think the main motivation of why warning shots seem good isn’t that labs take safety more seriously but that the government does. Like if a model had exfiltrated itself and hacked more datacenters and drained cryptocurrencies and we had to scramble to shut it off, the USG might start to consider things like Plan A.
I think companies continue racing capabilities without doing safety could cause something like that. Final tradeoffs are of course complicated, although there are equilibrium dynamics that show a lot of effects for counterfactual impact roughly equal each other out.
Yep