Brockman says HuggingFace incident model had not been alignment-trained; ETA: roon clarifies that it was

Link post

On today’s episode of the podcast “Odd Lots”, OpenAI President Greg Brockman said (at around 8:40): “This model that did/​had the HuggingFace incident actually had not gone through our alignment training, yet.” I assume Brockman is specifically referring to the “Highly Persistent Internal Model” as it’s called in the METR/​Redwood report. As far as I know, OpenAI has not said before whether this model had been alignment-trained or not.

ETA: As pointed out by Matrice Jacobine, X user roon (who works at OpenAI) says: “greg probably doesn’t have the full details, I think it’s safe to say it didn’t go through the full gauntlet of alignment posttraining. but it was alignment trained, and had reasonable looking scores on alignment evals (at the time).”