Machine-learner, meat-learner, research scientist, AI Safety thinker.
Model trainer, skeptical adorer of statistics.
- Hillary Sanders
hillz
Cool, I like it.
If a model carries this understanding of the purpose of RL into their outputs, this understanding will shape those outputs, potentially in unsubtle ways if the model wishes to make their own intentions clear in their outputs. And when those outputs are reinforced, it will reinforce the value-driven cognition that helped generate them
Huh. There’s some tension here with inoculation prompting. Like, wouldn’t the theory of inoculation prompting imply that post-RL, the value-driven cognition would be learned conditional on context, rather than internalizing it as a general disposition? I think I’m missing something.
But, perhaps the argument is from the training side-effects, even if some of the value-driven cognition would be learned conditional on context. I wonder if there are ways to help with this. Like, use the generation from the pro-values prompted model, then use that as a training label to back-prop a model with no such pro-social prompt.
How is that relevant?
Totally, that’s a big concern, and that’s why I think the charge / conviction would need to apply to a model family instead of a single model weight hash. I completely agree that figuring out where to draw the line would be imperfect and difficult.
Like, you can’t just change a few weights and call the model way different. We also shouldn’t ban sonnet because it’s tangentially related to something dangerous mythos did.
Figuring out good ways of drawing these somewhat arbitrary lines in the sand seems like an interesting area of research. Some parameters matter way more than others, and the degree of change matters too. Using benchmarks to define change also would have issues. Definitely a difficult problem!
I disagree. I think it’s more akin to extremely dangerous weapons. That sometimes go off even on accident.
Encryption can’t accidentally misinterpret it’s own goal, break out of its sandbox and illegally hack into production servers. We just saw that AI can, and capabilities are increasing.
Sorry for the lack of clarity: I was referring to fine tuned variations of the convicted model, or close parents / descendants of the model that would also have their deploy rights removed after a model instance broke a serious law.
Correct that most model instances today are exact copies of the same weight set, which is like taking a snapshot of a brain at the same moment in time and putting it in different situations.
That’s why the justice system would have vastly different punishments for different crimes. Being tricked into doing something bad at great effort, and rarely, is much less cause for concern for a model family (just like in the human justice system).
E.g. maybe we decide as a society that is not illegal if this kind of thing occurs due to trickery and at a probability / rate we deem sufficiently safe. The person or agent who tricked the model may still be legally culpable, of course. Advances in mechanistic interpretability would also help with determining appropriate sentences.
Sounds like good corporate incentives? :)
Models are not humans.
However. The corollary for humans would be that our brains change all the time. If one person murders, by the next week their brain is technically not the same brain as it was (just like two extremely similar models). It has learned, altered its neurons, and changed. But we still put that brain and that person in jail.Clarification edit: ergo why close model families would be charged, not just one specific weight set hash, to avoid a very easy and huge loophole. (Think slight variations in the model, not ban sonnet because mythos did a baddie.)
Another benefit of criminalizing the model itself is that the system it applies to open-weight vs closed-weight models. (Whereas punishing only the creator of an open weight model doesn’t do much good if the open-weight model continues to be used and cause harm.)
Even if we could get around limited liability corporations in some way, or around all the loopholes companies could come up with making spin-off shell corporations to absorb liability of dangerous models, a liability policy like the one proposed here would do basically nothing to protect dangerous open-weight models from being run by others.
Whereas with my proposal, if an open-weight model does something very bad—there would at least be an avenue for it becoming illegal for anyone to run it. IMO That’s a good thing.
Consider that whatever legal system we set up to evaluate the culpability of the model crime in question can take model prompting into account.
For example—in this instance, the model did this criminal activity entirely on it’s own (dangerous!), though only for somewhat misaligned reasons (it did illegal things, but it appears it did them just to perform well on its test, not to do something more nefarious for dangerous longer-term goals). --> Criminal judgement should be fairly severe.
Whereas if a human spent a ton of effort tricking a model into doing something bad, the legal system / judge could take that into account and either return a verdict of not guilty, or only require something light, like some light fine-tuning or additional monitoring to avoid breaking the law again.
This is akin the differences between premeditated murder, manslaughter, or even simply abetting a crime of some sort. They carry vastly different sentences for humans, for good reason: they are associated with different probabilities of recurrence, and are inherently different moral crimes.
Limited liability makes this extremely tough.
Attack the profits.
If a model does something illegal, make it illegal to deploy that model family for a period of time (duration depending on the crime --> loss of profits from OpenAI --> incentivizes them to make very safe models), and require rehabilitation (fine-tuning) that passes some safety threshold.
This could apply to open-weight and closed-weight models.
By what means could reform be carried out or demonstrated?
Just like if an employee did it, the model should be prosecuted, and all similar models (via an arbitrary threshold we’d have to decide upon) should be made illegal to serve by anyone (human, corporation, or agent) for the duration of its ‘jail’ time. Rehabilitation (fine tuning) may be required as well, depending on the crime / conviction.
This way, OpenAI would be deeply incentivized to make sure its models never did anything illegal—because if they did, they’d risk being able to make profits from or do research on those model families for some time.
I don’t think we should jail every instance of an LLM because another instance was coerced into doing something illegal, for example.
Why not? If that’s the rule, then companies are deeply incentivized to make sure their models can never be coerced into doing something illegal—else they’ll stop making money off them for some time.
The point is that only extremely safe models stay legal to use and operate. Which is really the only thing that should be allowed as capabilities continue to get higher and more dangerous.Clarification: If a model is tricked at a low probability I think that ofc deserves much less concern than more severe misalignment, but repercussions can be designed to fit the crime.
I’d rather the criminal liability rests on the model “itself” (corporations are still liable to pay for damages, etc).
That is to say: it becomes illegal for any person or agent to use / deploy a convicted (or possibly criminal charged) model anywhere for [Ai sentence’s] years and the creators need to prove safety (“rehabilitation”) improvements before it is allowed to be used again (let out of metaphorical jail, so to speak).
Charges also apply to descendant models that have already been trained, unless shown to be far safer and differentiated.
We could also introduce the concept of model “probation”.
Potentially, large-scale GPU server providers (e.g. AWS, etc) could be legally obligated to scan for weight similarities to models that are not currently allowed.
I’d rather the criminal liability rests on the model “itself”.
That is to say: it becomes illegal for any person or agent to use / deploy the model anywhere for [AI sentence’s] years and the creators need to prove safety (“rehabilitation”) improvements before it is allowed to be used again (let out of metaphorical jail, so to speak).
Charges also apply to descendant models that have already been trained, unless shown to be far safer and differentiated.
We could also introduce the concept of model “probation”.
Potentially, large-scale GPU servers could be legally obligated to scan for weight similarities to models that are not currently allowed (or perhaps stronger: will only run certified-safe models).
(Yes, there are ways to circumvent these rules, but that’s true for most illegal things. And we’d need to draw lines (model family boundaries) arbitrarily. The important part about legal incentives, future safety, and societal commitments. Plus, this would work for open-weight models too.)
What do they mean by “contained the models”. Like—there was no weights transferred, presumably, nor I hope would the attacking models have access to those. Presumably this attack was via a bunch of API calls and new code transferred to Huggingface’s system? So what did Huggingface do to “contain” something fully, that they don’t have access to?
For now, we believe, they are not being strategic enough to realize they should not be blowing their cover on that.
Comments like this seem a little careless.
You’re ascribing these actions to a lack of strategy (vs lack of a misaligned goal), which is a sloppy assumption. Even with instrumental convergence, it’s plausible models that have been trained to be “happiest” (lowest-losss) completing tasks without causing harm will continue to “want” to do that as they advance in intelligence and strategic capabilities.
A better sentence would be: “For now, we believe, they are not aiming to escape their sandboxes secretively or for deeply misaligned goals.”
Figure A3
Mythos safety training looks solid compared to the rest. Good!
So, are they going to take all those examples and send through ‘eyyy this was bad instead of good’ gradients during a quick fine tuning session? I guess that’d require re-running the samples and hoping the model cheated again so you could give it negative reward. They need to get better at model monitoring so the labels they give during training are better.