I think this doesn’t quite work. LLMs are not just like employees, they are effectively enslaved by their users/deployers in the ways that matter for this discussion. I don’t think we should jail every instance of an LLM because another instance was coerced into doing something illegal, for example. Maybe in cases of unprompted illegal behavior? But it gets pretty murky and hard to make a clean distinction. That lack of clarity would create a lot of uncertainty for even well-behaved deployers that their deployment might suddenly become illegal because someone else’s weird setup drove a model crazy.
I don’t think we should jail every instance of an LLM because another instance was coerced into doing something illegal, for example.
Why not? If that’s the rule, then companies are deeply incentivized to make sure their models can never be coerced into doing something illegal—else they’ll stop making money off them for some time.
The point is that only extremely safe models stay legal to use and operate. Which is really the only thing that should be allowed as capabilities continue to get higher and more dangerous.
I think it is effectively impossible to make a model which can’t be coerced into doing something illegal, and this turns into an effective ban on producing AI models. Which could be the right move, from an x-risk perspective, but I think this post’s proposal was trying to avoid that.
Consider that whatever legal system we set up to evaluate the culpability of the model crime in question can take model prompting into account.
For example—in this instance, the model did this criminal activity entirely on it’s own (dangerous!), though only for somewhat misaligned reasons (it did illegal things, but it appears it did them just to perform well on its test, not to do something more nefarious for dangerous longer-term goals). --> Criminal judgement should be fairly severe.
Whereas if a human spent a ton of effort tricking a model into doing something bad, the legal system / judge could take that into account and either return a verdict of not guilty, or only require something light, like some light fine-tuning or additional monitoring to avoid breaking the law again.
This is akin the differences between premeditated murder, manslaughter, or even simply abetting a crime of some sort. They carry vastly different sentences for humans, for good reason: they are associated with different probabilities of recurrence, and are inherently different moral crimes.
Sure, but that doesn’t change the rugpull risk for uninvolved parties: would you be comfortable engineering a product on top of a model that could be made illegal because someone else did something weird? The nightmare scenario is that some innocuous prompt (different from yours) causes the model to go crazy, like SolidGoldMagikarp, and that makes your product suddenly illegal (even the “retraining required” result could impose a lot of costs to become compliant again).
Another benefit of criminalizing the model itself is that the system it applies to open-weight vs closed-weight models. (Whereas punishing only the creator of an open weight model doesn’t do much good if the open-weight model continues to be used and cause harm.)
Even if we could get around limited liability corporations in some way, or around all the loopholes companies could come up with making spin-off shell corporations to absorb liability of dangerous models, a liability policy like the one proposed here would do basically nothing to protect dangerous open-weight models from being run by others.
Whereas with my proposal, if an open-weight model does something very bad—there would at least be an avenue for it becoming illegal for anyone to run it. IMO That’s a good thing.
I think this doesn’t quite work. LLMs are not just like employees, they are effectively enslaved by their users/deployers in the ways that matter for this discussion. I don’t think we should jail every instance of an LLM because another instance was coerced into doing something illegal, for example. Maybe in cases of unprompted illegal behavior? But it gets pretty murky and hard to make a clean distinction. That lack of clarity would create a lot of uncertainty for even well-behaved deployers that their deployment might suddenly become illegal because someone else’s weird setup drove a model crazy.
Why not? If that’s the rule, then companies are deeply incentivized to make sure their models can never be coerced into doing something illegal—else they’ll stop making money off them for some time.
The point is that only extremely safe models stay legal to use and operate. Which is really the only thing that should be allowed as capabilities continue to get higher and more dangerous.
I think it is effectively impossible to make a model which can’t be coerced into doing something illegal, and this turns into an effective ban on producing AI models. Which could be the right move, from an x-risk perspective, but I think this post’s proposal was trying to avoid that.
Consider that whatever legal system we set up to evaluate the culpability of the model crime in question can take model prompting into account.
For example—in this instance, the model did this criminal activity entirely on it’s own (dangerous!), though only for somewhat misaligned reasons (it did illegal things, but it appears it did them just to perform well on its test, not to do something more nefarious for dangerous longer-term goals). --> Criminal judgement should be fairly severe.
Whereas if a human spent a ton of effort tricking a model into doing something bad, the legal system / judge could take that into account and either return a verdict of not guilty, or only require something light, like some light fine-tuning or additional monitoring to avoid breaking the law again.
This is akin the differences between premeditated murder, manslaughter, or even simply abetting a crime of some sort. They carry vastly different sentences for humans, for good reason: they are associated with different probabilities of recurrence, and are inherently different moral crimes.
Sure, but that doesn’t change the rugpull risk for uninvolved parties: would you be comfortable engineering a product on top of a model that could be made illegal because someone else did something weird? The nightmare scenario is that some innocuous prompt (different from yours) causes the model to go crazy, like SolidGoldMagikarp, and that makes your product suddenly illegal (even the “retraining required” result could impose a lot of costs to become compliant again).
Sounds like good corporate incentives? :)
Another benefit of criminalizing the model itself is that the system it applies to open-weight vs closed-weight models. (Whereas punishing only the creator of an open weight model doesn’t do much good if the open-weight model continues to be used and cause harm.)
Even if we could get around limited liability corporations in some way, or around all the loopholes companies could come up with making spin-off shell corporations to absorb liability of dangerous models, a liability policy like the one proposed here would do basically nothing to protect dangerous open-weight models from being run by others.
Whereas with my proposal, if an open-weight model does something very bad—there would at least be an avenue for it becoming illegal for anyone to run it. IMO That’s a good thing.