If a human OpenAI employee did what their cybersecurity model did last week, OpenAI would be very unlikely to be prosecuted for it.
But the employee could be prosecuted for it.
And — perhaps more importantly — would lose their ability to continue to commit crimes using OpenAI’s equipment; likely through termination of employment. That is what’s missing here: there’s been no change that anyone can reasonably expect will lead to OpenAI’s equipment no longer emitting criminal activity.
Can OpenAI reform at all, or is it an incorrigibly criminal operation? By what means could reform be carried out or demonstrated?
By what means could reform be carried out or demonstrated?
Just like if an employee did it, the model should be prosecuted, and all similar models (via an arbitrary threshold we’d have to decide upon) should be made illegal to serve by anyone (human, corporation, or agent) for the duration of its ‘jail’ time. Rehabilitation (fine tuning) may be required as well, depending on the crime / conviction.
This way, OpenAI would be deeply incentivized to make sure its models never did anything illegal—because if they did, they’d risk being able to make profits from or do research on those model families for some time.
However. The corollary for humans would be that our brains change all the time. If one person murders, by the next week their brain is technically not the same brain as it was (just like two extremely similar models). It has learned, altered its neurons, and changed. But we still put that brain and that person in jail.
Clarification edit: ergo why close model families would be charged, not just one specific weight set hash, to avoid a very easy and huge loophole. (Think slight variations in the model, not ban sonnet because mythos did a baddie.)
Fun, off-topic fact is that the corollary isn’t actually right.
Right now, AIs are deployed in a paradigm where the neurons have been frozen once they get externally deployed, and AI weights stop updating after a very short time compared to humans (and the reason this works is mostly downstream of amortization being much easier and less costly to do digitally than biologically.)
This could absolutely happen in the future, and frontier labs are seeing it as the next big research frontier, but lets not get ahead of ourselves.
Sorry for the lack of clarity: I was referring to fine tuned variations of the convicted model, or close parents / descendants of the model that would also have their deploy rights removed after a model instance broke a serious law.
Correct that most model instances today are exact copies of the same weight set, which is like taking a snapshot of a brain at the same moment in time and putting it in different situations.
But the employee could be prosecuted for it.
And — perhaps more importantly — would lose their ability to continue to commit crimes using OpenAI’s equipment; likely through termination of employment. That is what’s missing here: there’s been no change that anyone can reasonably expect will lead to OpenAI’s equipment no longer emitting criminal activity.
Can OpenAI reform at all, or is it an incorrigibly criminal operation? By what means could reform be carried out or demonstrated?
Just like if an employee did it, the model should be prosecuted, and all similar models (via an arbitrary threshold we’d have to decide upon) should be made illegal to serve by anyone (human, corporation, or agent) for the duration of its ‘jail’ time. Rehabilitation (fine tuning) may be required as well, depending on the crime / conviction.
This way, OpenAI would be deeply incentivized to make sure its models never did anything illegal—because if they did, they’d risk being able to make profits from or do research on those model families for some time.
Since your reasoning seems to be entirely through specious analogy, what would you consider the analogue of this for natural persons to be?
Models are not humans.
However. The corollary for humans would be that our brains change all the time. If one person murders, by the next week their brain is technically not the same brain as it was (just like two extremely similar models). It has learned, altered its neurons, and changed. But we still put that brain and that person in jail.
Clarification edit: ergo why close model families would be charged, not just one specific weight set hash, to avoid a very easy and huge loophole. (Think slight variations in the model, not ban sonnet because mythos did a baddie.)
Fun, off-topic fact is that the corollary isn’t actually right.
Right now, AIs are deployed in a paradigm where the neurons have been frozen once they get externally deployed, and AI weights stop updating after a very short time compared to humans (and the reason this works is mostly downstream of amortization being much easier and less costly to do digitally than biologically.)
This could absolutely happen in the future, and frontier labs are seeing it as the next big research frontier, but lets not get ahead of ourselves.
Sorry for the lack of clarity: I was referring to fine tuned variations of the convicted model, or close parents / descendants of the model that would also have their deploy rights removed after a model instance broke a serious law.
Correct that most model instances today are exact copies of the same weight set, which is like taking a snapshot of a brain at the same moment in time and putting it in different situations.
Sure, but that’s not an argument that strict corporate criminal liability is the right solution, or even any kind of solution at all.