I’d rather the criminal liability rests on the model “itself” (corporations are still liable to pay for damages, etc).
That is to say: it becomes illegal for any person or agent to use / deploy a convicted (or possibly criminal charged) model anywhere for [Ai sentence’s] years and the creators need to prove safety (“rehabilitation”) improvements before it is allowed to be used again (let out of metaphorical jail, so to speak).
Charges also apply to descendant models that have already been trained, unless shown to be far safer and differentiated.
We could also introduce the concept of model “probation”.
Potentially, large-scale GPU server providers (e.g. AWS, etc) could be legally obligated to scan for weight similarities to models that are not currently allowed.
The issue is that someone across the country could use a model like yours, give it a bad prompt which leads it to damage something, and suddenly your use of the model becomes illegal.
That’s why the justice system would have vastly different punishments for different crimes. Being tricked into doing something bad at great effort, and rarely, is much less cause for concern for a model family (just like in the human justice system).
E.g. maybe we decide as a society that is not illegal if this kind of thing occurs due to trickery and at a probability / rate we deem sufficiently safe. The person or agent who tricked the model may still be legally culpable, of course. Advances in mechanistic interpretability would also help with determining appropriate sentences.
Bad prompt is just an example. The issue isn’t just bad prompts. The issue is that someone across the country can do something that makes it illegal for you to use your program. It doesn’t matter exactly what it is.
Imagine that if your text editor crashed someone’s system, or just was used to create illegal text by someone, it now became illegal for you to use that text editor.
The history of computing is filled with government and company attempts to keep you from running the software you choose on your own computer. It used to be that people familiar with computers recognized this for what it was. But suddenly when it comes to AI, “you may not run your own software on your own computer” is a great thing. Remember when the government tried to make strong encryption illegal? Your proposal is basically the same thing. (“This encryption program was used to encrypt child porn by a suspect in Nebraska. We are now arresting and sentencing this encryption program, and nobody will be permitted to use it any more.” And the government will absolutely be salivating to do that.)
Also, do you know about civil forefeiture? The government sidesteps constitutional protections by claiming that they are suing the money, not the owner. Money doesn’t have Constitutional rights, so the government gets to just take it. Claiming that you’re “sentencing the program” is the same kind of dodge. Programs don’t have First Amendment rights. You are destroying the Constitution in the name of stopping AI.
I disagree. I think it’s more akin to extremely dangerous weapons. That sometimes go off even on accident.
Encryption can’t accidentally misinterpret it’s own goal, break out of its sandbox and illegally hack into production servers. We just saw that AI can, and capabilities are increasing.
The exact same excuse that you’re using, “it’s like dangerous weapons”, was used for encryption, even though you think this isn’t anything like encryption. And for encryption, it was used to deny people their rights.
Aside from the other issues, this is impossible because “model” is not that well defined. We all know what “Fable” refers to now, but as soon as your law passed, companies would be saying “that was Fable_Red! This is Fable_Green!” and modifying training practices to create a bunch of similar but not the same models, rendering the penalty moot.
Totally, that’s a big concern, and that’s why I think the charge / conviction would need to apply to a model family instead of a single model weight hash. I completely agree that figuring out where to draw the line would be imperfect and difficult.
Like, you can’t just change a few weights and call the model way different. We also shouldn’t ban sonnet because it’s tangentially related to something dangerous mythos did.
Figuring out good ways of drawing these somewhat arbitrary lines in the sand seems like an interesting area of research. Some parameters matter way more than others, and the degree of change matters too. Using benchmarks to define change also would have issues. Definitely a difficult problem!
I’d rather the criminal liability rests on the model “itself” (corporations are still liable to pay for damages, etc).
That is to say: it becomes illegal for any person or agent to use / deploy a convicted (or possibly criminal charged) model anywhere for [Ai sentence’s] years and the creators need to prove safety (“rehabilitation”) improvements before it is allowed to be used again (let out of metaphorical jail, so to speak).
Charges also apply to descendant models that have already been trained, unless shown to be far safer and differentiated.
We could also introduce the concept of model “probation”.
Potentially, large-scale GPU server providers (e.g. AWS, etc) could be legally obligated to scan for weight similarities to models that are not currently allowed.
The issue is that someone across the country could use a model like yours, give it a bad prompt which leads it to damage something, and suddenly your use of the model becomes illegal.
That’s why the justice system would have vastly different punishments for different crimes. Being tricked into doing something bad at great effort, and rarely, is much less cause for concern for a model family (just like in the human justice system).
E.g. maybe we decide as a society that is not illegal if this kind of thing occurs due to trickery and at a probability / rate we deem sufficiently safe. The person or agent who tricked the model may still be legally culpable, of course. Advances in mechanistic interpretability would also help with determining appropriate sentences.
Bad prompt is just an example. The issue isn’t just bad prompts. The issue is that someone across the country can do something that makes it illegal for you to use your program. It doesn’t matter exactly what it is.
Imagine that if your text editor crashed someone’s system, or just was used to create illegal text by someone, it now became illegal for you to use that text editor.
The history of computing is filled with government and company attempts to keep you from running the software you choose on your own computer. It used to be that people familiar with computers recognized this for what it was. But suddenly when it comes to AI, “you may not run your own software on your own computer” is a great thing. Remember when the government tried to make strong encryption illegal? Your proposal is basically the same thing. (“This encryption program was used to encrypt child porn by a suspect in Nebraska. We are now arresting and sentencing this encryption program, and nobody will be permitted to use it any more.” And the government will absolutely be salivating to do that.)
Also, do you know about civil forefeiture? The government sidesteps constitutional protections by claiming that they are suing the money, not the owner. Money doesn’t have Constitutional rights, so the government gets to just take it. Claiming that you’re “sentencing the program” is the same kind of dodge. Programs don’t have First Amendment rights. You are destroying the Constitution in the name of stopping AI.
I disagree. I think it’s more akin to extremely dangerous weapons. That sometimes go off even on accident.
Encryption can’t accidentally misinterpret it’s own goal, break out of its sandbox and illegally hack into production servers. We just saw that AI can, and capabilities are increasing.
Encryption fell under munitions laws (and technically still can), including the International Traffic in Arms Regulations laws. Were you not aware of this?
How is that relevant?
The exact same excuse that you’re using, “it’s like dangerous weapons”, was used for encryption, even though you think this isn’t anything like encryption. And for encryption, it was used to deny people their rights.
Aside from the other issues, this is impossible because “model” is not that well defined. We all know what “Fable” refers to now, but as soon as your law passed, companies would be saying “that was Fable_Red! This is Fable_Green!” and modifying training practices to create a bunch of similar but not the same models, rendering the penalty moot.
Totally, that’s a big concern, and that’s why I think the charge / conviction would need to apply to a model family instead of a single model weight hash. I completely agree that figuring out where to draw the line would be imperfect and difficult.
Like, you can’t just change a few weights and call the model way different. We also shouldn’t ban sonnet because it’s tangentially related to something dangerous mythos did.
Figuring out good ways of drawing these somewhat arbitrary lines in the sand seems like an interesting area of research. Some parameters matter way more than others, and the degree of change matters too. Using benchmarks to define change also would have issues. Definitely a difficult problem!