The issue is that someone across the country could use a model like yours, give it a bad prompt which leads it to damage something, and suddenly your use of the model becomes illegal.
That’s why the justice system would have vastly different punishments for different crimes. Being tricked into doing something bad at great effort, and rarely, is much less cause for concern for a model family (just like in the human justice system).
E.g. maybe we decide as a society that is not illegal if this kind of thing occurs due to trickery and at a probability / rate we deem sufficiently safe. The person or agent who tricked the model may still be legally culpable, of course. Advances in mechanistic interpretability would also help with determining appropriate sentences.
The issue is that someone across the country could use a model like yours, give it a bad prompt which leads it to damage something, and suddenly your use of the model becomes illegal.
That’s why the justice system would have vastly different punishments for different crimes. Being tricked into doing something bad at great effort, and rarely, is much less cause for concern for a model family (just like in the human justice system).
E.g. maybe we decide as a society that is not illegal if this kind of thing occurs due to trickery and at a probability / rate we deem sufficiently safe. The person or agent who tricked the model may still be legally culpable, of course. Advances in mechanistic interpretability would also help with determining appropriate sentences.