Totally, that’s a big concern, and that’s why I think the charge / conviction would need to apply to a model family instead of a single model weight hash. I completely agree that figuring out where to draw the line would be imperfect and difficult.
Like, you can’t just change a few weights and call the model way different. We also shouldn’t ban sonnet because it’s tangentially related to something dangerous mythos did.
Figuring out good ways of drawing these somewhat arbitrary lines in the sand seems like an interesting area of research. Some parameters matter way more than others, and the degree of change matters too. Using benchmarks to define change also would have issues. Definitely a difficult problem!
Totally, that’s a big concern, and that’s why I think the charge / conviction would need to apply to a model family instead of a single model weight hash. I completely agree that figuring out where to draw the line would be imperfect and difficult.
Like, you can’t just change a few weights and call the model way different. We also shouldn’t ban sonnet because it’s tangentially related to something dangerous mythos did.
Figuring out good ways of drawing these somewhat arbitrary lines in the sand seems like an interesting area of research. Some parameters matter way more than others, and the degree of change matters too. Using benchmarks to define change also would have issues. Definitely a difficult problem!