It seems fairly easy to convince people to look into stuff like this even if you don’t have a lot of “social capital” to begin with. It seems pretty unusual for someone to outright deny things that there’s seemingly-damning evidence for. If you work at the AI company, know about the Hugging Face incident, etc. you probably have a reasonably high prior on AIs doing weird stuff like this, enough that you’d entertain the possibility that they could frame an employee. Companies probably don’t need a specific process that forces them to look into it when employees say they were framed, outside of whatever normal HR practices exist.
I suspect that maybe we have different levels of faith in regular HR operations than eachother?
HR as a discipline evolved to deal with very specific threat models related to legal liability, and maybe with the retention of useful talent against human actors. I don’t think current HR models (even in AI labs) are well equipped for something like this, and I suspect that it would not be taken very seriously as a possibility.
Okay, HR is probably willing to fire people even if they can’t prove without a shadow of a doubt that the misconduct happened, so I’m not necessarily comfortable depending on HR alone.
But if the employee reached out to the lab’s safety team, they’d probably be pretty willing to hear about this and conduct their own investigation if it sounds plausible. Hopefully there’s already a generic procedure for reaching out to the safety team about suspicious incidents across the company.
Probably the thing to do is just to raise awareness that this kind of thing could happen?
I think it’s probably good for something like this to happen by default, and not depend on social capital.
It seems fairly easy to convince people to look into stuff like this even if you don’t have a lot of “social capital” to begin with. It seems pretty unusual for someone to outright deny things that there’s seemingly-damning evidence for. If you work at the AI company, know about the Hugging Face incident, etc. you probably have a reasonably high prior on AIs doing weird stuff like this, enough that you’d entertain the possibility that they could frame an employee. Companies probably don’t need a specific process that forces them to look into it when employees say they were framed, outside of whatever normal HR practices exist.
I suspect that maybe we have different levels of faith in regular HR operations than eachother?
HR as a discipline evolved to deal with very specific threat models related to legal liability, and maybe with the retention of useful talent against human actors. I don’t think current HR models (even in AI labs) are well equipped for something like this, and I suspect that it would not be taken very seriously as a possibility.
Okay, HR is probably willing to fire people even if they can’t prove without a shadow of a doubt that the misconduct happened, so I’m not necessarily comfortable depending on HR alone.
But if the employee reached out to the lab’s safety team, they’d probably be pretty willing to hear about this and conduct their own investigation if it sounds plausible. Hopefully there’s already a generic procedure for reaching out to the safety team about suspicious incidents across the company.
Probably the thing to do is just to raise awareness that this kind of thing could happen?