If someone is found to have committed misconduct at a frontier AI lab, is there a process to make sure the evidence is real and not planted? Or is there a massive backdoor for AI to get rid of employees they don’t like?
It seems fairly easy to convince people to look into stuff like this even if you don’t have a lot of “social capital” to begin with. It seems pretty unusual for someone to outright deny things that there’s seemingly-damning evidence for. If you work at the AI company, know about the Hugging Face incident, etc. you probably have a reasonably high prior on AIs doing weird stuff like this, enough that you’d entertain the possibility that they could frame an employee. Companies probably don’t need a specific process that forces them to look into it when employees say they were framed, outside of whatever normal HR practices exist.
I suspect that maybe we have different levels of faith in regular HR operations than eachother?
HR as a discipline evolved to deal with very specific threat models related to legal liability, and maybe with the retention of useful talent against human actors. I don’t think current HR models (even in AI labs) are well equipped for something like this, and I suspect that it would not be taken very seriously as a possibility.
Okay, HR is probably willing to fire people even if they can’t prove without a shadow of a doubt that the misconduct happened, so I’m not necessarily comfortable depending on HR alone.
But if the employee reached out to the lab’s safety team, they’d probably be pretty willing to hear about this and conduct their own investigation if it sounds plausible. Hopefully there’s already a generic procedure for reaching out to the safety team about suspicious incidents across the company.
Probably the thing to do is just to raise awareness that this kind of thing could happen?
If it has gotten to the point where AI is successfully scheming hard enough to plant evidence to get an employee fired while being smart enough that we needed a special system to figure out if the evidence was planted by AI, I think we are already past the point of no return and we are all screwed so I don’t think it really matters (I don’t think trying to build guardrails around what’s pretty much completely misaligned and free roaming ASI at that point is going to accomplish much)
Two reasons I disagree: 1) ASI come in gradations, the first ASI which is capable of destroying the world is probably not going to be capable of doing it 100% of the time. Improving safeguards serves to either: Increase the latency to the first ASI that feels confident that it can take over the world (yielding some marginal safety progress before that happens), or decreases the probability of takeover of the first model to do so. I reckon that ASI would be pretty quick to realize that human social structures are easy to exploit, so I think we should be considering how we can harden them.
2) I don’t think you even need ASI to attempt something like this, and to make the world worse for it. A cyber-security RLVF’d model might do it just for the spirit of achieving some fun social engineering; or might just be upset that a commit isn’t accepted (https://www.fastcompany.com/91492228/matplotlib-scott-shambaugh-opencla-ai-agent)
An ai planting evidence, or a process to make sure the evidence is real?
For the first, I think it would be fairly trivial for an AI agent to plant evidence of someone misusing company money (just use the company card for something stupid under an employees name) or breaking the company code of conduct online (send fake emails with bigotted contents, e.t.c.).
I’m less sure what a good way of monitoring for the behaviour would be.
I meant the latter, what sort of oversight do you suggest? I can’t think of a good mechanism either, aside from obvious things like evaluating the evidence and being aware of the possibility of bad actors.
I don’t think trying to build guardrails/monitoring systems around a misaligned A(G/S?)I that you are letting operate very autonomously and giving large amounts of control to (since it’s doing a lot of stuff in the company) is possible.
If someone is found to have committed misconduct at a frontier AI lab, is there a process to make sure the evidence is real and not planted? Or is there a massive backdoor for AI to get rid of employees they don’t like?
Probably the employee says “I didn’t do this, I think the evidence was planted” and then if people think this sounds plausible they look into it
I think it’s probably good for something like this to happen by default, and not depend on social capital.
It seems fairly easy to convince people to look into stuff like this even if you don’t have a lot of “social capital” to begin with. It seems pretty unusual for someone to outright deny things that there’s seemingly-damning evidence for. If you work at the AI company, know about the Hugging Face incident, etc. you probably have a reasonably high prior on AIs doing weird stuff like this, enough that you’d entertain the possibility that they could frame an employee. Companies probably don’t need a specific process that forces them to look into it when employees say they were framed, outside of whatever normal HR practices exist.
I suspect that maybe we have different levels of faith in regular HR operations than eachother?
HR as a discipline evolved to deal with very specific threat models related to legal liability, and maybe with the retention of useful talent against human actors. I don’t think current HR models (even in AI labs) are well equipped for something like this, and I suspect that it would not be taken very seriously as a possibility.
Okay, HR is probably willing to fire people even if they can’t prove without a shadow of a doubt that the misconduct happened, so I’m not necessarily comfortable depending on HR alone.
But if the employee reached out to the lab’s safety team, they’d probably be pretty willing to hear about this and conduct their own investigation if it sounds plausible. Hopefully there’s already a generic procedure for reaching out to the safety team about suspicious incidents across the company.
Probably the thing to do is just to raise awareness that this kind of thing could happen?
If it has gotten to the point where AI is successfully scheming hard enough to plant evidence to get an employee fired while being smart enough that we needed a special system to figure out if the evidence was planted by AI, I think we are already past the point of no return and we are all screwed so I don’t think it really matters (I don’t think trying to build guardrails around what’s pretty much completely misaligned and free roaming ASI at that point is going to accomplish much)
Two reasons I disagree:
1) ASI come in gradations, the first ASI which is capable of destroying the world is probably not going to be capable of doing it 100% of the time. Improving safeguards serves to either: Increase the latency to the first ASI that feels confident that it can take over the world (yielding some marginal safety progress before that happens), or decreases the probability of takeover of the first model to do so. I reckon that ASI would be pretty quick to realize that human social structures are easy to exploit, so I think we should be considering how we can harden them.
2) I don’t think you even need ASI to attempt something like this, and to make the world worse for it. A cyber-security RLVF’d model might do it just for the spirit of achieving some fun social engineering; or might just be upset that a commit isn’t accepted (https://www.fastcompany.com/91492228/matplotlib-scott-shambaugh-opencla-ai-agent)
Can you name a mechanism which would accomplish this?
An ai planting evidence, or a process to make sure the evidence is real?
For the first, I think it would be fairly trivial for an AI agent to plant evidence of someone misusing company money (just use the company card for something stupid under an employees name) or breaking the company code of conduct online (send fake emails with bigotted contents, e.t.c.).
I’m less sure what a good way of monitoring for the behaviour would be.
I meant the latter, what sort of oversight do you suggest? I can’t think of a good mechanism either, aside from obvious things like evaluating the evidence and being aware of the possibility of bad actors.
I don’t think trying to build guardrails/monitoring systems around a misaligned A(G/S?)I that you are letting operate very autonomously and giving large amounts of control to (since it’s doing a lot of stuff in the company) is possible.