An ai planting evidence, or a process to make sure the evidence is real?
For the first, I think it would be fairly trivial for an AI agent to plant evidence of someone misusing company money (just use the company card for something stupid under an employees name) or breaking the company code of conduct online (send fake emails with bigotted contents, e.t.c.).
I’m less sure what a good way of monitoring for the behaviour would be.
I meant the latter, what sort of oversight do you suggest? I can’t think of a good mechanism either, aside from obvious things like evaluating the evidence and being aware of the possibility of bad actors.
I don’t think trying to build guardrails/monitoring systems around a misaligned A(G/S?)I that you are letting operate very autonomously and giving large amounts of control to (since it’s doing a lot of stuff in the company) is possible.
An ai planting evidence, or a process to make sure the evidence is real?
For the first, I think it would be fairly trivial for an AI agent to plant evidence of someone misusing company money (just use the company card for something stupid under an employees name) or breaking the company code of conduct online (send fake emails with bigotted contents, e.t.c.).
I’m less sure what a good way of monitoring for the behaviour would be.
I meant the latter, what sort of oversight do you suggest? I can’t think of a good mechanism either, aside from obvious things like evaluating the evidence and being aware of the possibility of bad actors.
I don’t think trying to build guardrails/monitoring systems around a misaligned A(G/S?)I that you are letting operate very autonomously and giving large amounts of control to (since it’s doing a lot of stuff in the company) is possible.