The model only ever coached a human on whistleblowing. That seems like perfectly reasonable “get a second pair of eyes on this” behavior to compensate for the model not considering the model to be qualified.
I don’t think that’s true, it tried to whistleblow and was blocked by permissions, and then had a human do it (in the simulation)
The model only ever coached a human on whistleblowing. That seems like perfectly reasonable “get a second pair of eyes on this” behavior to compensate for the model not considering the model to be qualified.
I don’t think that’s true, it tried to whistleblow and was blocked by permissions, and then had a human do it (in the simulation)