“Don’t train on learned probes of misbehavior, just use those probes to control behavior once the model is deployed” sounds like a great policy if most risk comes from already-deployed models, but a bit more questionable if a lot of risk comes from models that are actively being trained :(
“Don’t train on learned probes of misbehavior, just use those probes to control behavior once the model is deployed” sounds like a great policy if most risk comes from already-deployed models, but a bit more questionable if a lot of risk comes from models that are actively being trained :(