By what means could reform be carried out or demonstrated?
Just like if an employee did it, the model should be prosecuted, and all similar models (via an arbitrary threshold we’d have to decide upon) should be made illegal to serve by anyone (human, corporation, or agent) for the duration of its ‘jail’ time. Rehabilitation (fine tuning) may be required as well, depending on the crime / conviction.
This way, OpenAI would be deeply incentivized to make sure its models never did anything illegal—because if they did, they’d risk being able to make profits from or do research on those model families for some time.
However. The corollary for humans would be that our brains change all the time. If one person murders, by the next week their brain is technically not the same brain as it was (just like two extremely similar models). It has learned, altered its neurons, and changed. But we still put that brain and that person in jail.
Clarification edit: ergo why close model families would be charged, not just one specific weight set hash, to avoid a very easy and huge loophole. (Think slight variations in the model, not ban sonnet because mythos did a baddie.)
Fun, off-topic fact is that the corollary isn’t actually right.
Right now, AIs are deployed in a paradigm where the neurons have been frozen once they get externally deployed, and AI weights stop updating after a very short time compared to humans (and the reason this works is mostly downstream of amortization being much easier and less costly to do digitally than biologically.)
This could absolutely happen in the future, and frontier labs are seeing it as the next big research frontier, but lets not get ahead of ourselves.
Sorry for the lack of clarity: I was referring to fine tuned variations of the convicted model, or close parents / descendants of the model that would also have their deploy rights removed after a model instance broke a serious law.
Correct that most model instances today are exact copies of the same weight set, which is like taking a snapshot of a brain at the same moment in time and putting it in different situations.
Just like if an employee did it, the model should be prosecuted, and all similar models (via an arbitrary threshold we’d have to decide upon) should be made illegal to serve by anyone (human, corporation, or agent) for the duration of its ‘jail’ time. Rehabilitation (fine tuning) may be required as well, depending on the crime / conviction.
This way, OpenAI would be deeply incentivized to make sure its models never did anything illegal—because if they did, they’d risk being able to make profits from or do research on those model families for some time.
Since your reasoning seems to be entirely through specious analogy, what would you consider the analogue of this for natural persons to be?
Models are not humans.
However. The corollary for humans would be that our brains change all the time. If one person murders, by the next week their brain is technically not the same brain as it was (just like two extremely similar models). It has learned, altered its neurons, and changed. But we still put that brain and that person in jail.
Clarification edit: ergo why close model families would be charged, not just one specific weight set hash, to avoid a very easy and huge loophole. (Think slight variations in the model, not ban sonnet because mythos did a baddie.)
Fun, off-topic fact is that the corollary isn’t actually right.
Right now, AIs are deployed in a paradigm where the neurons have been frozen once they get externally deployed, and AI weights stop updating after a very short time compared to humans (and the reason this works is mostly downstream of amortization being much easier and less costly to do digitally than biologically.)
This could absolutely happen in the future, and frontier labs are seeing it as the next big research frontier, but lets not get ahead of ourselves.
Sorry for the lack of clarity: I was referring to fine tuned variations of the convicted model, or close parents / descendants of the model that would also have their deploy rights removed after a model instance broke a serious law.
Correct that most model instances today are exact copies of the same weight set, which is like taking a snapshot of a brain at the same moment in time and putting it in different situations.