In addition, strict liability is also commonly applied to owning farm animals. If they cause damage, the owner is responsible regardless of intent. It seems rather natural to extend this to AIs as well. I’d rather not get into evaluating the offending AIs intent, and would instead consider it’s actions as actions taken by the owner under strict liability rules.
In existing law, this generally applies to civil liability, not criminal. If my farm animals wander onto my neighbor’s property and cause damage, I am civilly responsible for damages. I am NOT criminally responsible as if I had trespassed and intentionally caused the damage myself.
This post is mostly reasonable if I read it as proposing a civil-liability standard. Unfortunately, it really sounds like it’s proposing a criminal-liability standard, which does not seem to me like a reasonable approach.
It is definitely reasonable to say that I am civilly liable for their medical bills.
It is plausibly reasonable to impose some pain-and-suffering damages, or to accuse me of negligence, especially if my dog has bitten people before.
It seems to me frankly deranged to say that this should be treated legally as if I myself had intentionally bitten the victim.
I appreciate that AI looks quite worrying, but I do not think that “okay, banning AI development directly looks like a hard political sell, but maybe we can find a sneaky way to pervert liability law and make it de facto impossible” is the sort of thought process that leads to good outcomes.
The aim here is not to make AI deployment de-facto impossible. In general, except in the most egregious cases, even when companies are found criminally liable the CEO is rarely sent to prison—instead the company is fined or otherwise punished.
Instead the aim is to incentivise AI companies to invest enough in safeguards that the level of fines is far lower than profit.
What about individual users of AI? It seems like kind of a cop-out to ask for criminal liability when it results in the same kinds of risks as civil liability for companies, but actual prison as a possibility for individuals.
What is the point of introducing criminal liability at all here, if most of the relevant actors are not individuals who can suffer actual criminal penalties, with leniency for individuals to compensate for the imbalance? Why not just stick with civil liability in the first place (genuinely asking, I am not a lawyer)?
The correct course of action for any individual wanting to deploy AI under this regime would be to create an LLC as a criminal liability condom solely for the AI deployment, which feels like a pointless bureaucratic hoop. If it is important to apply criminal law here but we don’t want to actually send individuals to prison, maybe we should treat AI deployments by individuals as if they were done in an LLC even if no LLC was created beforehand.
I’d rather the criminal liability rests on the model “itself”.
That is to say: it becomes illegal for any person or agent to use / deploy the model anywhere for [AI sentence’s] years and the creators need to prove safety (“rehabilitation”) improvements before it is allowed to be used again (let out of metaphorical jail, so to speak).
Charges also apply to descendant models that have already been trained, unless shown to be far safer and differentiated.
We could also introduce the concept of model “probation”.
Potentially, large-scale GPU servers could be legally obligated to scan for weight similarities to models that are not currently allowed.
(Yes, there are ways to circumvent these rules, but that’s true for most illegal things. And we’d need to draw lines (model family boundaries) arbitrarily. The important part about legal incentives, future safety, and societal commitments. Plus, this would work for open-weight models too.)
I think this doesn’t quite work. LLMs are not just like employees, they are effectively enslaved by their users/deployers in the ways that matter for this discussion. I don’t think we should jail every instance of an LLM because another instance was coerced into doing something illegal, for example. Maybe in cases of unprompted illegal behavior? But it gets pretty murky and hard to make a clean distinction. That lack of clarity would create a lot of uncertainty for even well-behaved deployers that their deployment might suddenly become illegal because someone else’s weird setup drove a model crazy.
I don’t think we should jail every instance of an LLM because another instance was coerced into doing something illegal, for example.
Why not? If that’s the rule, then companies are deeply incentivized to make sure their models can never be coerced into doing something illegal—else they’ll stop making money off them for some time.
The point is that only extremely safe models stay legal to use and operate. Which is really the only thing that should be allowed as capabilities continue to get higher and more dangerous.
I think it is effectively impossible to make a model which can’t be coerced into doing something illegal, and this turns into an effective ban on producing AI models. Which could be the right move, from an x-risk perspective, but I think this post’s proposal was trying to avoid that.
Consider that whatever legal system we set up to evaluate the culpability of the model crime in question can take model prompting into account.
For example—in this instance, the model did this criminal activity entirely on it’s own (dangerous!), though only for somewhat misaligned reasons (it did illegal things, but it appears it did them just to perform well on its test, not to do something more nefarious for dangerous longer-term goals). --> Criminal judgement should be fairly severe.
Whereas if a human spent a ton of effort tricking a model into doing something bad, the legal system / judge could take that into account and either return a verdict of not guilty, or only require something light, like some light fine-tuning or additional monitoring to avoid breaking the law again.
This is akin the differences between premeditated murder, manslaughter, or even simply abetting a crime of some sort. They carry vastly different sentences for humans, for good reason: they are associated with different probabilities of recurrence, and are inherently different moral crimes.
Sure, but that doesn’t change the rugpull risk for uninvolved parties: would you be comfortable engineering a product on top of a model that could be made illegal because someone else did something weird? The nightmare scenario is that some innocuous prompt (different from yours) causes the model to go crazy, like SolidGoldMagikarp, and that makes your product suddenly illegal (even the “retraining required” result could impose a lot of costs to become compliant again).
Another benefit of criminalizing the model itself is that the system it applies to open-weight vs closed-weight models. (Whereas punishing only the creator of an open weight model doesn’t do much good if the open-weight model continues to be used and cause harm.)
Even if we could get around limited liability corporations in some way, or around all the loopholes companies could come up with making spin-off shell corporations to absorb liability of dangerous models, a liability policy like the one proposed here would do basically nothing to protect dangerous open-weight models from being run by others.
Whereas with my proposal, if an open-weight model does something very bad—there would at least be an avenue for it becoming illegal for anyone to run it. IMO That’s a good thing.
In most of the cases that have been brought up, misconduct by an employee would not result in the criminal prosecution of a corporation either. As the page you linked explains, state law usually doesn’t impose criminal liability for one-off misconduct by rank-and-file employees, and while federal law might theoretically allow prosecution, doing so would in most cases be contrary to Justice Department guidelines that have a similar effect to the state laws.
If a human OpenAI employee did what their cybersecurity model did last week, OpenAI would be very unlikely to be prosecuted for it.
I don’t think the facts of any of the cases in the “deaths linked to chatbots” Wikipedia article would support a criminal prosecution even of an individual human. When the article mentions legal action taken in response, it’s either civil lawsuits or legislatures summoning AI lab executives to publicly yell at them.
If a human OpenAI employee did what their cybersecurity model did last week, OpenAI would be very unlikely to be prosecuted for it.
But the employee could be prosecuted for it.
And — perhaps more importantly — would lose their ability to continue to commit crimes using OpenAI’s equipment; likely through termination of employment. That is what’s missing here: there’s been no change that anyone can reasonably expect will lead to OpenAI’s equipment no longer emitting criminal activity.
Can OpenAI reform at all, or is it an incorrigibly criminal operation? By what means could reform be carried out or demonstrated?
By what means could reform be carried out or demonstrated?
Just like if an employee did it, the model should be prosecuted, and all similar models (via an arbitrary threshold we’d have to decide upon) should be made illegal to serve by anyone (human, corporation, or agent) for the duration of its ‘jail’ time. Rehabilitation (fine tuning) may be required as well, depending on the crime / conviction.
This way, OpenAI would be deeply incentivized to make sure its models never did anything illegal—because if they did, they’d risk being able to make profits from or do research on those model families for some time.
However. The corollary for humans would be that our brains change all the time. If one person murders, by the next week their brain is technically not the same brain as it was (just like two extremely similar models). It has learned, altered its neurons, and changed. But we still put that brain and that person in jail.
Clarification edit: ergo why close model families would be charged, not just one specific weight set hash, to avoid a very easy and huge loophole. (Think slight variations in the model, not ban sonnet because mythos did a baddie.)
Fun, off-topic fact is that the corollary isn’t actually right.
Right now, AIs are deployed in a paradigm where the neurons have been frozen once they get externally deployed, and AI weights stop updating after a very short time compared to humans (and the reason this works is mostly downstream of amortization being much easier and less costly to do digitally than biologically.)
This could absolutely happen in the future, and frontier labs are seeing it as the next big research frontier, but lets not get ahead of ourselves.
Sorry for the lack of clarity: I was referring to fine tuned variations of the convicted model, or close parents / descendants of the model that would also have their deploy rights removed after a model instance broke a serious law.
Correct that most model instances today are exact copies of the same weight set, which is like taking a snapshot of a brain at the same moment in time and putting it in different situations.
I agree these cases are not particularly problematic. This is preparation for worse cases, and also provides a standard which can be used to clarify existing cases so companies can proceed with confidence as to what they need to be worried about and what not.
Most of those don’t seem like they’d result in corporate criminal charges if a human employee did them either. Maybe the first one if the employee’s activities had a big enough impact on the corporation’s overall product roadmap or similar, but I would expect the prospect of a civil suit from the victim (which can already happen under existing law) to be a bigger deterrent to doing something risky than a highly uncertain possibility of criminal liability.
If a corporation screws up badly enough then authorities might try to throw the book at them by all available means, but that can already include criminal charges, presumably on some kind of theory of criminal negligence.
So I still don’t think you’ve given an example of a scenario where a model’s actions don’t presently expose its operator to criminal liability, but would under your proposal, such that that prospect of criminal liability could plausibly make the model not worth deploying when it otherwise would be.
The reason I bought up intent is because cyber security laws do actually depend on intent. We don’t prosecute somebody who accidentally triggers a remote execution exploit, but we do to somebody who did it on purpose.
We don’t want to rewrite the legal code for AI, so need to work out how to apply existing law to it.
Triggering a remote exection vulnerability accidentally is exceedingly unlikely to cause any serious damage anyway; that’ll just crash the process. Proper exploits do not happen accidentally. If the software has a bug that makes an accidental action cause damage then liability is (or at least should be) on whoever hosts or distributes that program.
In some other cases the intent might actually matter. It’ll require major rewriting of legal code anyway, if you want the intent of an AI to be something that can be considered here.
In addition, strict liability is also commonly applied to owning farm animals. If they cause damage, the owner is responsible regardless of intent. It seems rather natural to extend this to AIs as well. I’d rather not get into evaluating the offending AIs intent, and would instead consider it’s actions as actions taken by the owner under strict liability rules.
In existing law, this generally applies to civil liability, not criminal. If my farm animals wander onto my neighbor’s property and cause damage, I am civilly responsible for damages. I am NOT criminally responsible as if I had trespassed and intentionally caused the damage myself.
This post is mostly reasonable if I read it as proposing a civil-liability standard. Unfortunately, it really sounds like it’s proposing a criminal-liability standard, which does not seem to me like a reasonable approach.
I am indeed focused on criminal liability. Why doesn’t that seem to you as not a reasonable approach?
Say my dog bites someone, and they need stitches.
It is definitely reasonable to say that I am civilly liable for their medical bills.
It is plausibly reasonable to impose some pain-and-suffering damages, or to accuse me of negligence, especially if my dog has bitten people before.
It seems to me frankly deranged to say that this should be treated legally as if I myself had intentionally bitten the victim.
I appreciate that AI looks quite worrying, but I do not think that “okay, banning AI development directly looks like a hard political sell, but maybe we can find a sneaky way to pervert liability law and make it de facto impossible” is the sort of thought process that leads to good outcomes.
The aim here is not to make AI deployment de-facto impossible. In general, except in the most egregious cases, even when companies are found criminally liable the CEO is rarely sent to prison—instead the company is fined or otherwise punished.
Instead the aim is to incentivise AI companies to invest enough in safeguards that the level of fines is far lower than profit.
What about individual users of AI? It seems like kind of a cop-out to ask for criminal liability when it results in the same kinds of risks as civil liability for companies, but actual prison as a possibility for individuals.
This only applies if you deploy the AI yourself, so a tiny percentage of users (especially since most of the risk is from frontier models).
However I agree that we should be more lenient with them, at the point where we draft legislation we can haggle on the finer points.
What is the point of introducing criminal liability at all here, if most of the relevant actors are not individuals who can suffer actual criminal penalties, with leniency for individuals to compensate for the imbalance? Why not just stick with civil liability in the first place (genuinely asking, I am not a lawyer)?
The correct course of action for any individual wanting to deploy AI under this regime would be to create an LLC as a criminal liability condom solely for the AI deployment, which feels like a pointless bureaucratic hoop. If it is important to apply criminal law here but we don’t want to actually send individuals to prison, maybe we should treat AI deployments by individuals as if they were done in an LLC even if no LLC was created beforehand.
I’d rather the criminal liability rests on the model “itself”.
That is to say: it becomes illegal for any person or agent to use / deploy the model anywhere for [AI sentence’s] years and the creators need to prove safety (“rehabilitation”) improvements before it is allowed to be used again (let out of metaphorical jail, so to speak).
Charges also apply to descendant models that have already been trained, unless shown to be far safer and differentiated.
We could also introduce the concept of model “probation”.
Potentially, large-scale GPU servers could be legally obligated to scan for weight similarities to models that are not currently allowed.
(Yes, there are ways to circumvent these rules, but that’s true for most illegal things. And we’d need to draw lines (model family boundaries) arbitrarily. The important part about legal incentives, future safety, and societal commitments. Plus, this would work for open-weight models too.)
I think this doesn’t quite work. LLMs are not just like employees, they are effectively enslaved by their users/deployers in the ways that matter for this discussion. I don’t think we should jail every instance of an LLM because another instance was coerced into doing something illegal, for example. Maybe in cases of unprompted illegal behavior? But it gets pretty murky and hard to make a clean distinction. That lack of clarity would create a lot of uncertainty for even well-behaved deployers that their deployment might suddenly become illegal because someone else’s weird setup drove a model crazy.
Why not? If that’s the rule, then companies are deeply incentivized to make sure their models can never be coerced into doing something illegal—else they’ll stop making money off them for some time.
The point is that only extremely safe models stay legal to use and operate. Which is really the only thing that should be allowed as capabilities continue to get higher and more dangerous.
I think it is effectively impossible to make a model which can’t be coerced into doing something illegal, and this turns into an effective ban on producing AI models. Which could be the right move, from an x-risk perspective, but I think this post’s proposal was trying to avoid that.
Consider that whatever legal system we set up to evaluate the culpability of the model crime in question can take model prompting into account.
For example—in this instance, the model did this criminal activity entirely on it’s own (dangerous!), though only for somewhat misaligned reasons (it did illegal things, but it appears it did them just to perform well on its test, not to do something more nefarious for dangerous longer-term goals). --> Criminal judgement should be fairly severe.
Whereas if a human spent a ton of effort tricking a model into doing something bad, the legal system / judge could take that into account and either return a verdict of not guilty, or only require something light, like some light fine-tuning or additional monitoring to avoid breaking the law again.
This is akin the differences between premeditated murder, manslaughter, or even simply abetting a crime of some sort. They carry vastly different sentences for humans, for good reason: they are associated with different probabilities of recurrence, and are inherently different moral crimes.
Sure, but that doesn’t change the rugpull risk for uninvolved parties: would you be comfortable engineering a product on top of a model that could be made illegal because someone else did something weird? The nightmare scenario is that some innocuous prompt (different from yours) causes the model to go crazy, like SolidGoldMagikarp, and that makes your product suddenly illegal (even the “retraining required” result could impose a lot of costs to become compliant again).
Sounds like good corporate incentives? :)
Another benefit of criminalizing the model itself is that the system it applies to open-weight vs closed-weight models. (Whereas punishing only the creator of an open weight model doesn’t do much good if the open-weight model continues to be used and cause harm.)
Even if we could get around limited liability corporations in some way, or around all the loopholes companies could come up with making spin-off shell corporations to absorb liability of dangerous models, a liability policy like the one proposed here would do basically nothing to protect dangerous open-weight models from being run by others.
Whereas with my proposal, if an open-weight model does something very bad—there would at least be an avenue for it becoming illegal for anyone to run it. IMO That’s a good thing.
The reason is that we have existing criminal law and want to apply it to what the AI agent does.
And yes, that proposal of treating an individual as if he was the employer of the LLM is a reasonable approach.
My argument is AI is far more similar to an employee than a dog.
In most of the cases that have been brought up, misconduct by an employee would not result in the criminal prosecution of a corporation either. As the page you linked explains, state law usually doesn’t impose criminal liability for one-off misconduct by rank-and-file employees, and while federal law might theoretically allow prosecution, doing so would in most cases be contrary to Justice Department guidelines that have a similar effect to the state laws.
If a human OpenAI employee did what their cybersecurity model did last week, OpenAI would be very unlikely to be prosecuted for it.
I don’t think the facts of any of the cases in the “deaths linked to chatbots” Wikipedia article would support a criminal prosecution even of an individual human. When the article mentions legal action taken in response, it’s either civil lawsuits or legislatures summoning AI lab executives to publicly yell at them.
Are there other kinds of cases you had in mind?
But the employee could be prosecuted for it.
And — perhaps more importantly — would lose their ability to continue to commit crimes using OpenAI’s equipment; likely through termination of employment. That is what’s missing here: there’s been no change that anyone can reasonably expect will lead to OpenAI’s equipment no longer emitting criminal activity.
Can OpenAI reform at all, or is it an incorrigibly criminal operation? By what means could reform be carried out or demonstrated?
Just like if an employee did it, the model should be prosecuted, and all similar models (via an arbitrary threshold we’d have to decide upon) should be made illegal to serve by anyone (human, corporation, or agent) for the duration of its ‘jail’ time. Rehabilitation (fine tuning) may be required as well, depending on the crime / conviction.
This way, OpenAI would be deeply incentivized to make sure its models never did anything illegal—because if they did, they’d risk being able to make profits from or do research on those model families for some time.
Since your reasoning seems to be entirely through specious analogy, what would you consider the analogue of this for natural persons to be?
Models are not humans.
However. The corollary for humans would be that our brains change all the time. If one person murders, by the next week their brain is technically not the same brain as it was (just like two extremely similar models). It has learned, altered its neurons, and changed. But we still put that brain and that person in jail.
Clarification edit: ergo why close model families would be charged, not just one specific weight set hash, to avoid a very easy and huge loophole. (Think slight variations in the model, not ban sonnet because mythos did a baddie.)
Fun, off-topic fact is that the corollary isn’t actually right.
Right now, AIs are deployed in a paradigm where the neurons have been frozen once they get externally deployed, and AI weights stop updating after a very short time compared to humans (and the reason this works is mostly downstream of amortization being much easier and less costly to do digitally than biologically.)
This could absolutely happen in the future, and frontier labs are seeing it as the next big research frontier, but lets not get ahead of ourselves.
Sorry for the lack of clarity: I was referring to fine tuned variations of the convicted model, or close parents / descendants of the model that would also have their deploy rights removed after a model instance broke a serious law.
Correct that most model instances today are exact copies of the same weight set, which is like taking a snapshot of a brain at the same moment in time and putting it in different situations.
Sure, but that’s not an argument that strict corporate criminal liability is the right solution, or even any kind of solution at all.
I agree these cases are not particularly problematic. This is preparation for worse cases, and also provides a standard which can be used to clarify existing cases so companies can proceed with confidence as to what they need to be worried about and what not.
Can you please give an example of a case where you think a no-fault criminal liability standard for AI would both be helpful and make legal sense?
If a model was asked to research a topic and stole the results from a competitor.
If a model gave concrete advice about how to carry out a terrorist attack.
If a model agreed to take control of a car and crashed it into someone.
If...
Most of those don’t seem like they’d result in corporate criminal charges if a human employee did them either. Maybe the first one if the employee’s activities had a big enough impact on the corporation’s overall product roadmap or similar, but I would expect the prospect of a civil suit from the victim (which can already happen under existing law) to be a bigger deterrent to doing something risky than a highly uncertain possibility of criminal liability.
If a corporation screws up badly enough then authorities might try to throw the book at them by all available means, but that can already include criminal charges, presumably on some kind of theory of criminal negligence.
So I still don’t think you’ve given an example of a scenario where a model’s actions don’t presently expose its operator to criminal liability, but would under your proposal, such that that prospect of criminal liability could plausibly make the model not worth deploying when it otherwise would be.
They almost definitely would prosecute the company if this became a regular pattern (and not just a one off).
The reason I bought up intent is because cyber security laws do actually depend on intent. We don’t prosecute somebody who accidentally triggers a remote execution exploit, but we do to somebody who did it on purpose.
We don’t want to rewrite the legal code for AI, so need to work out how to apply existing law to it.
Triggering a remote exection vulnerability accidentally is exceedingly unlikely to cause any serious damage anyway; that’ll just crash the process. Proper exploits do not happen accidentally. If the software has a bug that makes an accidental action cause damage then liability is (or at least should be) on whoever hosts or distributes that program.
In some other cases the intent might actually matter. It’ll require major rewriting of legal code anyway, if you want the intent of an AI to be something that can be considered here.
But an AI isn’t an animal, it is a machine, and we do not apply strict liability to machines that malfunction.