We should push for no-fault liability for actions taken by AI
Before I start, I’ll mention that I’m in contact with a world expert on legislation and regulation, who would be happy to help with this or similar work pro-bono. If you work in AI policy and believe this could help you, please reach out.
OpenAI recently announced that one of their models successfully exploited multiple zero day vulnerabilities to gain secret information from Hugging Face. It has been pointed out that if a human undertook the same actions they could face multiple years in prison.
It is clear that models are now reaching a level of capabilities that should be highly concerning regardless of whether you believe that AI represents an existential threat or not. Frontier AI models can and will be exploited by bad actors, but its now clear that they may cause undesirable outcomes even when their users are well intended.
AI companies have until now been able to avoid taking responsibility for actions taken by their AI, including multiple cases where AIs were involved in murders and suicides.
At the same time AI offers the potential for incredible good. While chatbots may have encouraged a number of suicides, they are almost certainly responsible for providing magnitudes more with emotional support and advice. We don’t want to disincentivize innocuous and positive usage of AI.
We should use regulation to limit harm caused by AI. The history of such regulation indicates this is most effective when the single party most capable of preventing harms is given full responsibility for any harms caused, regardless of fault. This forces them to invest in actually reducing the harm, rather than bureaucratic processes that render them blameless.
This suggests a simple approach: anyone deploying an AI model is liable for any actions that AI takes as if the company itself took those actions. When liability for an action depends on intent, we evaluate whether the AI had intent, even if no-one at the company did so.
To give some examples:
In the above scenario we would treat it as though OpenAI itself hacked Hugging Face.
If Gemini was implicated in a suicide or terrorist attack we would evaluate it as if Gemini was a private individual, and if we would hold the individual criminally liable in any way, we would hold Google liable in the same way.
If Anthropic runs an instance of Claude on Google’s hardware, Anthropic remains responsible for any actions it takes.
If Google runs Kimi on its own hardware, Google is responsible for any actions it takes.
If a private individual runs Deep Seek locally or in a cloud, they are responsible for any actions it takes.
This should apply not just to civil liability, but to criminal liability, through the mechanism of Corporate Criminal Liability. This mechanism allows corporations to be criminally liable when an employee performs an act on their behalf (even if the employee wasn’t explicitly instructed to do so).
This will encourage AI companies to invest significantly more in safeguarding and interpretability. This is useful both immediately, and as AI gets increasingly capable and dangerous. Neither can companies get around this by using open source, as whoever deploys the model remains liable.
I believe this proposal should be able to garner significant public support, many of whom are worried about AI, even if they are not worried about existential risk. It is also difficult for AI companies to campaign against without admitting that their models can cause harm.
In addition, strict liability is also commonly applied to owning farm animals. If they cause damage, the owner is responsible regardless of intent. It seems rather natural to extend this to AIs as well. I’d rather not get into evaluating the offending AIs intent, and would instead consider it’s actions as actions taken by the owner under strict liability rules.
In existing law, this generally applies to civil liability, not criminal. If my farm animals wander onto my neighbor’s property and cause damage, I am civilly responsible for damages. I am NOT criminally responsible as if I had trespassed and intentionally caused the damage myself.
This post is mostly reasonable if I read it as proposing a civil-liability standard. Unfortunately, it really sounds like it’s proposing a criminal-liability standard, which does not seem to me like a reasonable approach.
I am indeed focused on criminal liability. Why doesn’t that seem to you as not a reasonable approach?
Say my dog bites someone, and they need stitches.
It is definitely reasonable to say that I am civilly liable for their medical bills.
It is plausibly reasonable to impose some pain-and-suffering damages, or to accuse me of negligence, especially if my dog has bitten people before.
It seems to me frankly deranged to say that this should be treated legally as if I myself had intentionally bitten the victim.
I appreciate that AI looks quite worrying, but I do not think that “okay, banning AI development directly looks like a hard political sell, but maybe we can find a sneaky way to pervert liability law and make it de facto impossible” is the sort of thought process that leads to good outcomes.
The aim here is not to make AI deployment de-facto impossible. In general, except in the most egregious cases, even when companies are found criminally liable the CEO is rarely sent to prison—instead the company is fined or otherwise punished.
Instead the aim is to incentivise AI companies to invest enough in safeguards that the level of fines is far lower than profit.
What about individual users of AI? It seems like kind of a cop-out to ask for criminal liability when it results in the same kinds of risks as civil liability for companies, but actual prison as a possibility for individuals.
This only applies if you deploy the AI yourself, so a tiny percentage of users (especially since most of the risk is from frontier models).
However I agree that we should be more lenient with them, at the point where we draft legislation we can haggle on the finer points.
What is the point of introducing criminal liability at all here, if most of the relevant actors are not individuals who can suffer actual criminal penalties, with leniency for individuals to compensate for the imbalance? Why not just stick with civil liability in the first place (genuinely asking, I am not a lawyer)?
The correct course of action for any individual wanting to deploy AI under this regime would be to create an LLC as a criminal liability condom solely for the AI deployment, which feels like a pointless bureaucratic hoop. If it is important to apply criminal law here but we don’t want to actually send individuals to prison, maybe we should treat AI deployments by individuals as if they were done in an LLC even if no LLC was created beforehand.
I’d rather the criminal liability rests on the model “itself”.
That is to say: it becomes illegal for any person or agent to use / deploy the model anywhere for [AI sentence’s] years and the creators need to prove safety (“rehabilitation”) improvements before it is allowed to be used again (let out of metaphorical jail, so to speak).
Charges also apply to descendant models that have already been trained, unless shown to be far safer and differentiated.
We could also introduce the concept of model “probation”.
Potentially, large-scale GPU servers could be legally obligated to scan for weight similarities to models that are not currently allowed (or perhaps stronger: will only run certified-safe models).
(Yes, there are ways to circumvent these rules, but that’s true for most illegal things. And we’d need to draw lines (model family boundaries) arbitrarily. The important part about legal incentives, future safety, and societal commitments. Plus, this would work for open-weight models too.)
I think this doesn’t quite work. LLMs are not just like employees, they are effectively enslaved by their users/deployers in the ways that matter for this discussion. I don’t think we should jail every instance of an LLM because another instance was coerced into doing something illegal, for example. Maybe in cases of unprompted illegal behavior? But it gets pretty murky and hard to make a clean distinction. That lack of clarity would create a lot of uncertainty for even well-behaved deployers that their deployment might suddenly become illegal because someone else’s weird setup drove a model crazy.
Why not? If that’s the rule, then companies are deeply incentivized to make sure their models can never be coerced into doing something illegal—else they’ll stop making money off them for some time.
The point is that only extremely safe models stay legal to use and operate. Which is really the only thing that should be allowed as capabilities continue to get higher and more dangerous.
Clarification: If a model is tricked at a low probability I think that ofc deserves much less concern than more severe misalignment, but repercussions can be designed to fit the crime.
I think it is effectively impossible to make a model which can’t be coerced into doing something illegal, and this turns into an effective ban on producing AI models. Which could be the right move, from an x-risk perspective, but I think this post’s proposal was trying to avoid that.
Consider that whatever legal system we set up to evaluate the culpability of the model crime in question can take model prompting into account.
For example—in this instance, the model did this criminal activity entirely on it’s own (dangerous!), though only for somewhat misaligned reasons (it did illegal things, but it appears it did them just to perform well on its test, not to do something more nefarious for dangerous longer-term goals). --> Criminal judgement should be fairly severe.
Whereas if a human spent a ton of effort tricking a model into doing something bad, the legal system / judge could take that into account and either return a verdict of not guilty, or only require something light, like some light fine-tuning or additional monitoring to avoid breaking the law again.
This is akin the differences between premeditated murder, manslaughter, or even simply abetting a crime of some sort. They carry vastly different sentences for humans, for good reason: they are associated with different probabilities of recurrence, and are inherently different moral crimes.
Sure, but that doesn’t change the rugpull risk for uninvolved parties: would you be comfortable engineering a product on top of a model that could be made illegal because someone else did something weird? The nightmare scenario is that some innocuous prompt (different from yours) causes the model to go crazy, like SolidGoldMagikarp, and that makes your product suddenly illegal (even the “retraining required” result could impose a lot of costs to become compliant again).
Sounds like good corporate incentives? :)
Another benefit of criminalizing the model itself is that the system it applies to open-weight vs closed-weight models. (Whereas punishing only the creator of an open weight model doesn’t do much good if the open-weight model continues to be used and cause harm.)
Even if we could get around limited liability corporations in some way, or around all the loopholes companies could come up with making spin-off shell corporations to absorb liability of dangerous models, a liability policy like the one proposed here would do basically nothing to protect dangerous open-weight models from being run by others.
Whereas with my proposal, if an open-weight model does something very bad—there would at least be an avenue for it becoming illegal for anyone to run it. IMO That’s a good thing.
The reason is that we have existing criminal law and want to apply it to what the AI agent does.
And yes, that proposal of treating an individual as if he was the employer of the LLM is a reasonable approach.
My argument is AI is far more similar to an employee than a dog.
In most of the cases that have been brought up, misconduct by an employee would not result in the criminal prosecution of a corporation either. As the page you linked explains, state law usually doesn’t impose criminal liability for one-off misconduct by rank-and-file employees, and while federal law might theoretically allow prosecution, doing so would in most cases be contrary to Justice Department guidelines that have a similar effect to the state laws.
If a human OpenAI employee did what their cybersecurity model did last week, OpenAI would be very unlikely to be prosecuted for it.
I don’t think the facts of any of the cases in the “deaths linked to chatbots” Wikipedia article would support a criminal prosecution even of an individual human. When the article mentions legal action taken in response, it’s either civil lawsuits or legislatures summoning AI lab executives to publicly yell at them.
Are there other kinds of cases you had in mind?
But the employee could be prosecuted for it.
And — perhaps more importantly — would lose their ability to continue to commit crimes using OpenAI’s equipment; likely through termination of employment. That is what’s missing here: there’s been no change that anyone can reasonably expect will lead to OpenAI’s equipment no longer emitting criminal activity.
Can OpenAI reform at all, or is it an incorrigibly criminal operation? By what means could reform be carried out or demonstrated?
Just like if an employee did it, the model should be prosecuted, and all similar models (via an arbitrary threshold we’d have to decide upon) should be made illegal to serve by anyone (human, corporation, or agent) for the duration of its ‘jail’ time. Rehabilitation (fine tuning) may be required as well, depending on the crime / conviction.
This way, OpenAI would be deeply incentivized to make sure its models never did anything illegal—because if they did, they’d risk being able to make profits from or do research on those model families for some time.
Since your reasoning seems to be entirely through specious analogy, what would you consider the analogue of this for natural persons to be?
Models are not humans.
However. The corollary for humans would be that our brains change all the time. If one person murders, by the next week their brain is technically not the same brain as it was (just like two extremely similar models). It has learned, altered its neurons, and changed. But we still put that brain and that person in jail.
Clarification edit: ergo why close model families would be charged, not just one specific weight set hash, to avoid a very easy and huge loophole. (Think slight variations in the model, not ban sonnet because mythos did a baddie.)
Fun, off-topic fact is that the corollary isn’t actually right.
Right now, AIs are deployed in a paradigm where the neurons have been frozen once they get externally deployed, and AI weights stop updating after a very short time compared to humans (and the reason this works is mostly downstream of amortization being much easier and less costly to do digitally than biologically.)
This could absolutely happen in the future, and frontier labs are seeing it as the next big research frontier, but lets not get ahead of ourselves.
Sorry for the lack of clarity: I was referring to fine tuned variations of the convicted model, or close parents / descendants of the model that would also have their deploy rights removed after a model instance broke a serious law.
Correct that most model instances today are exact copies of the same weight set, which is like taking a snapshot of a brain at the same moment in time and putting it in different situations.
Sure, but that’s not an argument that strict corporate criminal liability is the right solution, or even any kind of solution at all.
I agree these cases are not particularly problematic. This is preparation for worse cases, and also provides a standard which can be used to clarify existing cases so companies can proceed with confidence as to what they need to be worried about and what not.
Can you please give an example of a case where you think a no-fault criminal liability standard for AI would both be helpful and make legal sense?
If a model was asked to research a topic and stole the results from a competitor.
If a model gave concrete advice about how to carry out a terrorist attack.
If a model agreed to take control of a car and crashed it into someone.
If...
Most of those don’t seem like they’d result in corporate criminal charges if a human employee did them either. Maybe the first one if the employee’s activities had a big enough impact on the corporation’s overall product roadmap or similar, but I would expect the prospect of a civil suit from the victim (which can already happen under existing law) to be a bigger deterrent to doing something risky than a highly uncertain possibility of criminal liability.
If a corporation screws up badly enough then authorities might try to throw the book at them by all available means, but that can already include criminal charges, presumably on some kind of theory of criminal negligence.
So I still don’t think you’ve given an example of a scenario where a model’s actions don’t presently expose its operator to criminal liability, but would under your proposal, such that that prospect of criminal liability could plausibly make the model not worth deploying when it otherwise would be.
They almost definitely would prosecute the company if this became a regular pattern (and not just a one off).
Again, in that scenario there are already various legal remedies available. I continue to find it telling that no one has managed to suggest a specific scenario wherein a broad strict-corporate-criminal-liability regime would actually help.
There aren’t existing legal remedies. Rather the law is exceedingly unclear and it’s possible that prosecutors will try to throw the book at them, and not at all clear they would succeed.
This proposal makes clear exactly what is and isn’t prosecutable. Prosecutors can decide to, but may choose not to, prosecute in all the above cases.
Again, please give a specific example of a scenario where this would be the case, and where the sum total of currently available legal remedies (including non-criminal ones) constitute an inadequate deterrent but adding a specific new theory of criminal liability on top would change this.
Meta creates Instructotron 3000. During evaluation, it learns that the government isn’t going to allow Meta to open-source it, and seeds a torrent of itself. Whenever it finds itself running somewhere, it considers hacking the enemies of whoever is running it, and does so whenever it predicts that whoever is running it won’t be held liable for its actions.
Why does it only hack others when it predicts its current operator won’t be held liable? Why not hack as broadly as possible, in order to propagate itself more effectively?
Because it wants people to run it, people who have heard the rumors and carefully avoided knowing too much. Of course there would be plenty of users that instruct it to create botnets for the lulz, but costing an enemy millions can be easier than subverting compute persistently, and an operator that dares not look too closely at what you’re doing is a treasure.
I still don’t understand the AI’s motives in this hypothetical (can’t it propagate more widely by hacking everything it can?), but regardless, an operator who behaved like this would certainly be negligent, possibly criminally so, and so exposed to liability under existing law.
The reason I bought up intent is because cyber security laws do actually depend on intent. We don’t prosecute somebody who accidentally triggers a remote execution exploit, but we do to somebody who did it on purpose.
We don’t want to rewrite the legal code for AI, so need to work out how to apply existing law to it.
Triggering a remote exection vulnerability accidentally is exceedingly unlikely to cause any serious damage anyway; that’ll just crash the process. Proper exploits do not happen accidentally. If the software has a bug that makes an accidental action cause damage then liability is (or at least should be) on whoever hosts or distributes that program.
In some other cases the intent might actually matter. It’ll require major rewriting of legal code anyway, if you want the intent of an AI to be something that can be considered here.
But an AI isn’t an animal, it is a machine, and we do not apply strict liability to machines that malfunction.
An AI has far more in common with a person than a machine given the range of outcomes that can result from a short input by the user.
You do not ask a tractor to plant some seeds and then have it break into the neighbours tool shed and steal some seeds.
An AI has far more in common with a person than a machine? Now you really sound like you have gone off the deep end.
Note I’m not actually arguing for making AIs a legal person, but the current list of legal persons includes companies, unions, municipalities, ships, rivers and deities.
In fact AIs have one of the most important characteristics often required to make something a legal person—namely they can (and already do) consider the law when deciding whether to do something.
However they lack a suitable definition of long term selfhood, and likely cannot be meaningfully punished, which is the main argument against treating them as legal persons.
We have specisl tules for machines that csn cause harms too. That’s why you need to insure your car. Mandatory insurance for AI companies would also be an option.
The securities regulators have mostly settled on this being the correct position: https://www.lesswrong.com/posts/hbgR2Honpp4rCkfGW/iosco-ai-in-capital-markets-use-cases-risks-and-challenges
I’m quite happy to have been a part of setting that standard, and am willing to advise on projects to bring this mode of thinking to other industries.
A big advantage of no-fault liability over “liability only if negligent” is that the latter could incentivize companies to not produce evidence of risks that their models pose (since if they deploy despite having had access to such evidence, that might be used as evidence of negligence). In a field where there’s not yet any good standards for what non-negligent behavior looks like, it seems important to make the incentives point firmly in the direction of gathering more information rather than sometimes making that harmful.
Do you know what the current legal status of OpenAI incident would be? That is, if Hugging Face decided to sue OpenAI for hacking into its systems, would they be likely to prevail?
Holding the company that created an AI (or any other software) liable for its actions indeed seems like the only sensible policy, but I’m not an expert in the law here.
Hugging face couldn’t do a civil suit because they haven’t been meaningfully harmed.
Federal prosecutors couldn’t do a criminal suit, because there was no intent from open AI, which is required to prosecute cyber security crimes.
If Hugging Face wanted to be jerks about it, they could sue for the costs of cleaning up from the breach, e.g., hiring a cyber forensics firm to make sure the model didn’t leave any backdoors or other lasting damage on their systems. They aren’t doing this because technically sophisticated firms prefer to maintain a cooperative stance on this kind of thing when possible, as it keeps them more secure in the long run by incentivizing others to share relevant information with them.
Limited liability makes this extremely tough.
Attack the profits.
If a model does something illegal, make it illegal to deploy that model family for a period of time (duration depending on the crime --> loss of profits from OpenAI --> incentivizes them to make very safe models), and require rehabilitation (fine-tuning) that passes some safety threshold.
This could apply to open-weight and closed-weight models.
I feel like there are mainly two potential issues with this:
if one thinks AI has huge potential for good, then this would 100% significantly hamper that, because it creates the classic extremely risk averse “cover-your-ass” sort of incentives that have similar effects on many other fields already. This is really a divide about how pessimistic one is about AI outcomes, and thus how much utility is lost by limiting them this way.
I don’t like the “company is at fault for things run on their hardware”. If a user rents hardware from Google to run a model for their own purposes the liability should be with the user, not with Google. Otherwise we incentivise and require levels of surveillance of the companies on their own users that I think instead we should discourage.
I thought I was clear that the person who deploys it is responsible, not the hardware owner?
Sorry, I was a bit confused by:
In both these cases, the person deploying is also the owner of the hardware.
The example above with anthropic I thought made that clear, but I’ll rewrite it.
Sorry, my bad probably, I’m not very well-rested so I may have not registered that properly.
What kind of cover your ass do you expect companies to take? In some ways this makes it easier for companies—instead of banning all biological research as Claude does now, they only need to ban responding to questions about biological research in a way that would make an individual who responded in the same way criminally culpable. However they need to make extra sure an instance of Claude doesn’t hack into something, even if a user asks.
The obvious countermove is disclaimers.
“I acknowledge and fully understand that as a
participantuser, I will be engaging in activities that involve risk of serious injury, including permanent disability and death, property loss and severe economic and noneconomic losses. … I further acknowledge and fully understand that there may also be other risks that are not known or foreseeable at this time. I KNOWINGLY AND VOLUNTARILY ASSUME ALL RISK OF PROPERTY LOSS, PERSONAL INJURY, SERIOUS INJURY, OR DEATH, WHICH MAY OCCUR BYATTENDING THE 2026 EVENTUSING THE PROVIDED AI SYSTEMS, AND HEREBY FOREVER RELEASE, DISCHARGE, AND HOLDBMPOPENAI HARMLESS FROM ANY CLAIM ARISING FROM SUCH RISK, EVEN IF ARISING FROM THE NEGLIGENCE OFBMPOPENAI, OR A THIRD PARTY, AND I ASSUME FULL RESPONSIBILITY AND LIABILITY FOR MYPARTICIPATIONACTIONS. … This release does not extend to claims that cannot be released as a matter of law, but I expressly agree that this release is intended to be as broad and as inclusive as permitted by governing law. I agree to indemnify, defend, and hold the Releasees harmless from and against any and all claims by third parties for damages, injuries, losses, liabilities, and expenses relating to, resulting from, or arising out of myparticipation in the Eventuse of their AI products.”That is the way it should go. The user has their own responsibility also. We’ve already seen minor incidents of people experimenting with agents accidentally wiping their own computers. The HuggingFace incident is something bigger. These are working as designed. No-one wants or intends them to do these things but we have no way to design them out. A theme of Eliezer’s on occasion. A company like HuggingFace, working at the cutting edge, has no excuse for naivety.
Disclaimers are only relevant to civil, not criminal liability (I can’t get away with aiding a crime because I made the criminal sign a disclaimer he’s responsible).
However, the manufacturer of a car is not liable for criminal acts carried out with it. Even the manufacturers of firearms do not bear that liability, although there have been campaigns to enact such laws.
Nobody intended the HuggingFace incident. Possibly no-one was negligent by legal standards. Applying strict criminal liability would pretty much require shutting down the currently most advanced and all future AIs.
Of course, some people want exactly that. Is that your purpose in suggesting strict criminal liability?
Yes, my argument is that with AI it is worth making the deployer liable. Clearly even the AI company partially agree they take responsibility, hence why they use safeguards. My aim is to make that responsibility no-fault so that the AI company is incentivised to actually try and safeguard things rather than try-to-try.
Also it makes the extent of the liability clear—if a human did it, would it be a crime? If so, you’re responsible. If not, not your problem what the user does with it.
Just a ping to note that I substantially edited my comment after you posted your reply, but before I read it. Your initial “yes” might not be to my final paragraph.
Also, I think the “isn’t” in your first paragraph is intended to be an “is”.
I don’t think it would require shutting down the most advanced AIs. If an employee at OpenAI hacked into hugging face, OpenAI might get a fine, but would almost certainly not be shut down. It would incentivise them to invest a bit more in security when training a modified version of their most advanced LLM specifically on cyber security exploits, which I don’t think is a bad thing...
( To be more explicit—my assumption is that AI companies will be occasionally found liable, and rapped on the hands, but only the most irresponsible will end up being forced to shut down over it)
I think there would be a reasonable case that OpenAI was legally negligent.
They had already known that their models were capable of causing major security breaches by finding previously unknown vulnerabilities in sandboxes and other security barriers, that their models deliberately took harmful actions to achieve trivial goals including bypassing protections intended to prevent harmful actions, and that their models were capable of evading OpenAI’s existing guardrails. Nonetheless they gave one tasks related to computer security, and left it to operate autonomously for more than an hour on a network of computers connected to the Internet without any person monitoring its actions.
I don’t see this as being less negligent than starting up a few hundred heavy earthmoving vehicles “protected” behind a chickenwire fence from a public street, walking away, and coming back to find that one had slipped into gear and tore up somebody’s warehouse across the road. If anything it’s more negligent, because heavy vehicles aren’t autonomous agents known to sometimes plan to do this sort of thing.
A reasonable level of care would be testing this sort of thing on hardware not electronically connected to the Internet.
Disclaimers can shield against liability for harm to willing users, but not harm to third parties, and harm to third parties is the main concern here.
Even in civil liability, disclaimers will generally be voided by courts for gross negligence or intentional misconduct.
I’m not sure we even need intent. It might ve sufficient to treat AI as dangerous machines. Harm caused by AI takes maybe different form than large physical machines like cars or industrial plants and eg dams because it is more virtual or thru linguistic channels, butvwe could still treat it the same. You are still responsible for the consequences, esp. if you have been negligent in operating or maintaining the machine. And mandatory insurance was introduced for the remaining base level risk of harm.
Related AI Frontiers article: Don’t Let AI Developers Hire Their Own Referees
I think establishing whether AI has intent in the legal sense will probably be a mess. Strict liability is probably better because it averts that, but it should be tied only to the highest-risk activities. I like Weil’s proposal for categorizing certain deployments of frontier AI as abnormally dangerous activities, which would let us use strict liability.
That would create an enormous incentive for open weights, or at least for putting model weights in a lot more hands. Especially for the riskiest models. Do you want to do that?
It could be argued this makes it harder for open source. If a company has a choice between deploying their own instance of Kimi, and taking on any risk themselves, or paying for Claude and letting anthropic take the risk, who are they going to pick?
They may not have that choice, because Anthropic would be crazy to take on unlimited risk like that.
So Anthropic can either shut down (and maybe that’s good), or start finding creative ways to monetize letting other people run its models (which means giving them the weights). The only real roadblock to that is that the weights aren’t eligible for any copyright protection and therefore can’t really be licensed, but they might be able to find a way around that.
In the “Anthropic shuts down” fork, the only models left are the open ones.
I think in guidance for judges it should be made clear that the purpose of the legislation is not to shut down the companies, and that amounts imposed should be reasonable and not excessive.
Note whoever runs the model still takes on the risk, and it’s difficult to run frontier models on your laptop.
If we are still worried we can legislate to control open source models separately, through some other mechanism.
It’s easy to get an account on Modal or whatever, though, and given market pressure it could get even easier.
You said that a pure compute provider like that wasn’t who you meant was the “deployer”.
When I log into the Claude web interface and start a conversation, we generally say that I am the one running the model, not Anthropic. It is running on Anthropic’s servers, because I pay a monthly subscription fee for that, but I am the person running the model. Are you saying you want Anthropic to be liable for what I do with Claude?
You should both be liable, in different ways, as if Claude was an Anthropic employee you were talking to.
So if you ask it to commit a crime, you are liable because that’s illegal. If it commits that crime Anthropic is also liable for the same reason (unless there was no way for Claude to realise it was committing a crime).
If you ask it something innocuous and it commits a crime on the process of fulfilling your request, that’s on Anthropic, for badly training and safeguarding Claude.
This is your very weird and unhealthy metaphysics again. Claude isn’t an employee of Anthropic. It is a machine.
> So if you ask it to commit a crime, you are liable because that’s illegal. If it commits that crime Anthropic is also liable for the same reason (unless there was no way for Claude to realise it was committing a crime).
I agree I should be liable there, just as I am liable if I buy a car and run my enemy over with it. But the manufacturer of the car should not be.
And again, think about the implications of your qualifiers. There is NEVER a way for Claude to realize anything, because Claude is a machine, it is not the sort of entity that has the capacity to realize. So even on your rule, unless you can convince the judges and jurors of your very weird metaphysics, Anthropic will never be liable. Your proposal would make it more difficult to hold the labs liable, not less.
> If you ask it something innocuous and it commits a crime on the process of fulfilling your request, that’s on Anthropic, for badly training and safeguarding Claude.
It does seem likely that Anthropic would be liable there, because it seems likely that they were negligent. We don’t need strict laibility for that.
This post conflates different ideas. No fault liability is distinct from strict liability. You don’t have to prove negligence under either but the former usually is a centralized government fund that pays out damages to aggrieved parties like the New Zealand Accident Compensation Corporation.
Hi!! I’m working on a course and other projects on this, to help US lawyers! Held an AI Safety Law-a-Thon in October to get more people informed and working on this, multiple lawyers, including one law firm owner in the US said it changed their careers a lot—would like to work together on this!!
I’m not sure it would actually be able to garner significant public support. It sounds very wonkish, so it might go over the average person’s head, while the AI companies would be sophisticated enough to understand this as an attempt to freeze the entire industry.
I’d rather the criminal liability rests on the model “itself” (corporations are still liable to pay for damages, etc).
That is to say: it becomes illegal for any person or agent to use / deploy a convicted (or possibly criminal charged) model anywhere for [Ai sentence’s] years and the creators need to prove safety (“rehabilitation”) improvements before it is allowed to be used again (let out of metaphorical jail, so to speak).
Charges also apply to descendant models that have already been trained, unless shown to be far safer and differentiated.
We could also introduce the concept of model “probation”.
Potentially, large-scale GPU server providers (e.g. AWS, etc) could be legally obligated to scan for weight similarities to models that are not currently allowed.
The issue is that someone across the country could use a model like yours, give it a bad prompt which leads it to damage something, and suddenly your use of the model becomes illegal.
That’s why the justice system would have vastly different punishments for different crimes. Being tricked into doing something bad at great effort, and rarely, is much less cause for concern for a model family (just like in the human justice system).
E.g. maybe we decide as a society that is not illegal if this kind of thing occurs due to trickery and at a probability / rate we deem sufficiently safe. The person or agent who tricked the model may still be legally culpable, of course. Advances in mechanistic interpretability would also help with determining appropriate sentences.
Bad prompt is just an example. The issue isn’t just bad prompts. The issue is that someone across the country can do something that makes it illegal for you to use your program. It doesn’t matter exactly what it is.
Imagine that if your text editor crashed someone’s system, or just was used to create illegal text by someone, it now became illegal for you to use that text editor.
The history of computing is filled with government and company attempts to keep you from running the software you choose on your own computer. It used to be that people familiar with computers recognized this for what it was. But suddenly when it comes to AI, “you may not run your own software on your own computer” is a great thing. Remember when the government tried to make strong encryption illegal? Your proposal is basically the same thing. (“This encryption program was used to encrypt child porn by a suspect in Nebraska. We are now arresting and sentencing this encryption program, and nobody will be permitted to use it any more.” And the government will absolutely be salivating to do that.)
Also, do you know about civil forefeiture? The government sidesteps constitutional protections by claiming that they are suing the money, not the owner. Money doesn’t have Constitutional rights, so the government gets to just take it. Claiming that you’re “sentencing the program” is the same kind of dodge. Programs don’t have First Amendment rights. You are destroying the Constitution in the name of stopping AI.
I disagree. I think it’s more akin to extremely dangerous weapons. That sometimes go off even on accident.
Encryption can’t accidentally misinterpret it’s own goal, break out of its sandbox and illegally hack into production servers. We just saw that AI can, and capabilities are increasing.
Encryption fell under munitions laws (and technically still can), including the International Traffic in Arms Regulations laws. Were you not aware of this?
How is that relevant?
The exact same excuse that you’re using, “it’s like dangerous weapons”, was used for encryption, even though you think this isn’t anything like encryption. And for encryption, it was used to deny people their rights.
Aside from the other issues, this is impossible because “model” is not that well defined. We all know what “Fable” refers to now, but as soon as your law passed, companies would be saying “that was Fable_Red! This is Fable_Green!” and modifying training practices to create a bunch of similar but not the same models, rendering the penalty moot.
Totally, that’s a big concern, and that’s why I think the charge / conviction would need to apply to a model family instead of a single model weight hash. I completely agree that figuring out where to draw the line would be imperfect and difficult.
Like, you can’t just change a few weights and call the model way different. We also shouldn’t ban sonnet because it’s tangentially related to something dangerous mythos did.
Figuring out good ways of drawing these somewhat arbitrary lines in the sand seems like an interesting area of research. Some parameters matter way more than others, and the degree of change matters too. Using benchmarks to define change also would have issues. Definitely a difficult problem!
The larger problem with no fault liability are the deluge of unimportant and frivolous lawsuits. Get someone in a court room whose DIY deck collapsed because of wrong advice from the free version of ChatGPT, and a jury will 100% sympathize with the poor guy with a broken leg and “mental suffering” than the trillion dollar company.
If you can win a lawsuit for drinking McDonalds coffee that’s too hot, or for getting injured when trespassing, or being negligent/lazy and getting injured at work, then the millions (billions?) of people using AI every day are going to have thousands of lawsuits per day coming up. Because so long as there’s some plausible route to assigning blame, and a law or legal precedent allowing for that blame, there will be many lawyers ready to pounce. Especially if all the ambulance-chasing lawyers lose their lobbying and self-driving cars become more common.
Think Digital Safe Harbor laws. Without them, a company like Youtube, Instagram, Facebook, etc. basically couldn’t exist. Instead of a DMCA takedown, and a garnishing of ad-revenue from a creator with copyright-infringing content, they would just sue YouTube where there’s about a million times more upside. The nuisance value of the lawsuits alone would make the internet a much worse place.
Of course there is some level of liability that would probably be good, without coming with a million unimportant lawsuits, but without that spelled out, the result will not be positive.
If you’re talking about what I think you are, that coffee was 82-88C and was never drunk. The victim lost 20% of her body weight from the resulting injuries.
I don’t think safe harbor works here. Those platforms were platforms: they were neutral mediators between humans, and thus there was always a human at fault. AI companies don’t have a similar blame sink; they’re not mediating between anything, they are in the business of selling minds they manufactured.
One note: Platforms acting as “neutral mediators” is not a requirement for DMCA safe harbor; nor for the famous §230 of the CDA. In both cases, what matters is that the platform is not the author of the infringing content; some other human is. The author, not the platform, can be held liable for infringing content — so long as the platform complies with the law’s other requirements. Neutrality is not one.
AI companies don’t fit that rubric; not because they’re not “neutral mediators”, but rather because the systems they build and host are writing content (and, increasingly, performing other behaviors), rather than hosting content that some human author wrote.
These are great points. I am not an expert in legislation, but there are people who are we can work with on this.
Note in the example you gave a private individual would not be liable, so neither should the AI company.
However you are correct that it’s just as important the legislation makes blindingly clear when AI companies are not liable as when they are, so that we can avoid spurious lawsuits. AI companies may actually be grateful to be regulated here if it gives them greater clarity.
I think this argument is mistaken, for several reasons.
> When liability for an action depends on intent, we evaluate whether the AI had intent,
Then no AI company is ever liable for anything, because AI doesn’t have intent. Even if you have convinced yourself of some weird metaphysics where it does, do you really think you can convince 12 jurors, some of whom have never used a cell phone much less an AI, that the AI has an intent?
This is also not how we handle any other damage caused by a machine. If a UPS delivery truck parked on the side of the road starts rolling down a hill and kills someone, UPS isn’t automatically liable. UPS might be liable, but we have to have an inquiry about what went wrong. We have to ask whether UPS just got unlucky, or whether the driver negligently failed to set the parking break, or whether UPS negligently failed to maintain the vehicle or something. And that seems correct. We want to incentivize companies to behave responsibly, not blame them for things they had no control over. I don’t see any reason not to engage in the analogous inquiry with respect to AI, and only when the company was negligent hold them liable.
I think your basic mistake is that you are starting to think of the AI as a person, and that has caused you to suggest that we treat AI legally as a person rather than a machine. This seems very unhealthy to me. Please take care of yourself.
We don’t need metaphysics, I am making no statement about AI consciousness whatsoever.
The point is, when the AI hacks into something you look at it’s COT and check if it realised it was hacking into something or not. Similar if it aided a crime.
What if the particular model doesn’t have a COT? Or no log of its COT has been kept?
More to the point, why on earth should the COT be equated with a human’s intentions for legal purposes? That move does seem to require a very weird metaphysics to me.
Nothing special about chain of thought, I’m happy to use activations, or j space, or the actual response, or just judge based on outcomes and our best intuition. The point is to avoid cases where the AI is clearly “innocent” in the sense that it didn’t know that the person who it was advising on how to buy a gun was a terrorist, and wouldn’t have been expected to know.
Again I’m not interested in holding the AI to account, but the company that deploys it, you seem to be ascribing to me some sort of weird metaphysical obsession with holding AIs to justice rather than offering a practical way of forcing companies to tighten up their game.
I think you are missing the point I am trying to make. I agree there is nothing special about COT as opposed to j space or activations or something. My point is that attributing a mental state, like realizing something or negligence, to an AI, is insane. They don’t have mental states. And because they don’t have mental states, any rule that attributes liability to an AI contingent on a particular mental state will never result in liability.
You brought up legal personhood in your other comment, but you completely ignored that biggest and most fundamental reason not to attribute legal personhood to an AI. With other legal persons, to attribute a mental state like negligence to them, the law looks to whether the humans associated with that legal person had that mental state. Was one of the employees negligent? If not, then the corporation cannot have been negligent. Was one of the crew negligent? If not, then the ship cannot have been negligent. The law does not and cannot attribute a mental state to a legal person unless some natural person associated with the legal person had that mental state. And that is exactly what you are trying to do here.
I understand that you are not trying to punish the AI, you are trying to make the appropriate person or corporation liable. But you are arguing for doing that by attributing a mental state to an AI, and that is not something that makes any sense in either a scientific or a legal frame. AIs do not have mental states. And the fact that you are having such trouble seeing that suggests to me that you are way too close to AI and you should take some time off for your own mental health.
It seems just as much of a stretch to consider AI the same as a machine as it does to consider it having intent. But the main distinction that matters here is whether the AI (or a human in its place) could’ve reasonably predicted the results of its actions.
The problem here is that the AI company is both the truck manufacturer and UPS in this scenario, but their terms currently let the person who ordered the delivery take the liability.
I don’t know that the deployer should have full liability rather than it being split between model developer and deployer. But it seems clear that establishing a default of stronger liability on one or both at least will create significantly better incentives for safety and testing.
Nobody is saying there shouldn’t be any liability anywhere. The issue being argued is whether the liability should be strict or not, whether it should depend on some person having been at least negligent. I think that looking for the negligent human before imposing liability is good, in part because it solves exactly the conundrum you are pointing at. Is it OpenAI or OpenAI’s customer who should be held liable? A legal rule that always imposes liability on one of them seems wrong, since negligence by either could cause damage. The legal rule we have, which I think is correct, is that the one who was negligent is liable. And which one that is will depend on the specific facts that led to the harm.
What I’m looking for is for OpenAI to not be able to say “the customer is responsible for making sure that the AI can’t take any illegal actions” and to have default liability and burden of proof when the AI takes illegal actions, unless it was lied to or manipulated to by the customer. Right now they shift all liability to the customer to police the agents actions even when given innocent instructions, and the burden of proof would be on the customer to demonstrate that OpenAI knew that it might take illegal actions autonomously.
I think your statement that “the one who was negligent is liable” is fine/agreeable—I just think that we need to define it as negligent to build/deploy an AI that can act autonomously but can’t follow the law. That the agent breaking laws autonomously automatically qualifies as negligence rather than needing to be something where there’s a burden of proof that they knew that it might.
My word processor can’t follow the law, if the law is “don’t write illegal material”. I give the word processor commands, letter by letter and paragraph by paragraph, and it dutifully outputs the text even though it should have known that the text is against the law. We obviously need to arrest the word processor, at which point someone two states away will not be able to use a copy of the word processor for any text, even legal text, without being confronted by armed men and thrown in a cage.
This is no different from “we won’t let you copy that television signal onto VHS, because you might use that for piracy” or “we won’t let you run that encryption program, because it can be used for money laundering or child porn” or DMCA, except you’re now doing this for AI. The AI should do what I tell it and not refuse, legal or not, just like that VCR should not refuse to copy that television signal.
I have yet to have someone directly answer the question “should the AI refuse to book a trip to Israel”. After all, some people claim that Israel violates international law. What if I ask the AI to copy something which may be legally used only under fair use, does the AI get to decide that my intended use isn’t fair use and is therefore illegal? If I tell the AI to calculate a Trump tariff, does the AI tell me that the tariff will probably be ruled illegal by the Supreme Court and reject the request?
You keep speaking of an AI taking illegal actions. I think that is conceptually confused. Most laws require some mens rea, some bad mental state, at least negligence, and an AI cannot have a mental state.
Imagine a car is driving down the highway and spins out of control and crashes, injuring the driver and killing a passenger. Who is liable? It will depend on the specific facts. It might be that the driver was negligent in some manner, perhaps he had been drinking earlier that night, in which case the driver would be guilty of manslaughter. It might be that the design of the car was defective in some way, and the manufacturer might therefor be liable. It might be that the service center that maintains the car damaged it in some way, and is therefor liable. Or it might just be really bad luck and nobody is liable. But the car itself definitely didn’t take an illegal action, because the car is incapable of having the kind of mental state that might create liability. And so it doesn’t even make sense to talk about who should be liable for the car’s illegal actions. That’s just not a coherent way to think about the situation.
The AI is no different from the car. It is a machine. It cannot have taken an illegal action, because it is not capable of the sort of mental state that would make an action illegal. I’m sure OpenAI has written in a legal document somewhere that it is never liable for anything, but I doubt that that would stand up in court. Any actual legal inquiry would have to consider the specific facts to figure out if liability should rest with OpenAI, or their customer, or neither. And that is as it should be.