A recurring sub-theme across several of my research interests this year have been various forms of deception checking, particularly automated deception checking.
I’ve gotten pretty disappointed in the space. Not all the time (eg Pangram is great), but consistently they can be bad, and bad in ways that are not obvious to outsiders or low-information buyers.
I believe I’ve identified a structural reason for this. If you’re a deception checking company, there’s a consistent tradeoff for what you can invest your resources in:
You can invest in better deception checking
You can invest in better deception. Specifically, you can invest in more and more elaborate lies about how your product totally works.
Across the board[1], it seems like many companies (perhaps correctly?) decided that the profit-maximizing move is #2.
This doesn’t work forever—eventually people wise up and are suspicious of the, ahem, AI Snake Oil that the deception detectors sell. And in fields where actual detectors do work (say Pangram for AI text detection), I think they eventually rise above the noise. This is probably not a field that you can keep lying forever, particularly when better alternative exist. But the lying lie detectors and the scamming scam detectors can keep lying and scamming for a long time. Every new form of deception can create a secondary grift window.
The existence proof and commonality of this dynamic so far should make us be suspicious of and guard against this dynamic continuing to happen. Especially as we enter new domains in AI epistemics, and the need for novel forms of deception detection.
Consider the first wave of superhumanly enhanced persuasive text and videos in the future. Afterwards, we might see an overflow of “detector” companies for superhuman manipulation, that don’t work but will try to persuade low-information buyers that they totally do work (possibly with superhumanly enhanced arguments in their own favor).
In the long run humanity can probably figure out which detectors actually work vs are fake, but also in the long run, we’re all...
AI text detection, Ai video (deepfake) detection, pre-2022 plagiarism detection, fraud recovery, human lie detection/polygraphs (which I hope to write about someday), etc.
I also think this is a really important issue, and that’s why I’m raising money for my startup to build an ai-powered fake ai detector detector… ah nevermind
I think some of this is also incentives from the customer side. Some number of customers want to be able to say things aren’t AI, without actually caring about the ground truth. Eg if you pay copywriters by volume, they have a strong incentive to use AI. And, if you don’t start checking until it’s ubiquitous, you risk losing your whole labor force once you do start checking, and the lower quality services are often cheaper, to boot.
Nobody actually says ‘I want to pay a third party to lie to my customers about my labor practices’; the CEO is just disincentivized from looking at the options on the market closely enough to tell the difference, and orders that the company go with the cheaper option.
I agree that helping to serve customer needs to outsource accountability/blame is one reason. I find it more understandable/almost morally gray when it comes to hiring corporate copywriters, and more unambiguously evil when it comes to, say, financial fraud deepfakes or polygraphs in criminal justice. Alas
A recurring sub-theme across several of my research interests this year have been various forms of deception checking, particularly automated deception checking.
I’ve gotten pretty disappointed in the space. Not all the time (eg Pangram is great), but consistently they can be bad, and bad in ways that are not obvious to outsiders or low-information buyers.
I believe I’ve identified a structural reason for this. If you’re a deception checking company, there’s a consistent tradeoff for what you can invest your resources in:
You can invest in better deception checking
You can invest in better deception. Specifically, you can invest in more and more elaborate lies about how your product totally works.
Across the board[1], it seems like many companies (perhaps correctly?) decided that the profit-maximizing move is #2.
This doesn’t work forever—eventually people wise up and are suspicious of the, ahem, AI Snake Oil that the deception detectors sell. And in fields where actual detectors do work (say Pangram for AI text detection), I think they eventually rise above the noise. This is probably not a field that you can keep lying forever, particularly when better alternative exist. But the lying lie detectors and the scamming scam detectors can keep lying and scamming for a long time. Every new form of deception can create a secondary grift window.
The existence proof and commonality of this dynamic so far should make us be suspicious of and guard against this dynamic continuing to happen. Especially as we enter new domains in AI epistemics, and the need for novel forms of deception detection.
Consider the first wave of superhumanly enhanced persuasive text and videos in the future. Afterwards, we might see an overflow of “detector” companies for superhuman manipulation, that don’t work but will try to persuade low-information buyers that they totally do work (possibly with superhumanly enhanced arguments in their own favor).
In the long run humanity can probably figure out which detectors actually work vs are fake, but also in the long run, we’re all...
AI text detection, Ai video (deepfake) detection, pre-2022 plagiarism detection, fraud recovery, human lie detection/polygraphs (which I hope to write about someday), etc.
For what it’s worth, Pangram also seems to only work on lazy prompts. I wasn’t even trying to avoid it and got a 100% human score on an AI-written post.
Yeah, they bias heavily towards avoiding false positives at the cost of avoiding false negatives!
I also think this is a really important issue, and that’s why I’m raising money for my startup to build an ai-powered fake ai detector detector… ah nevermind
I think some of this is also incentives from the customer side. Some number of customers want to be able to say things aren’t AI, without actually caring about the ground truth. Eg if you pay copywriters by volume, they have a strong incentive to use AI. And, if you don’t start checking until it’s ubiquitous, you risk losing your whole labor force once you do start checking, and the lower quality services are often cheaper, to boot.
Nobody actually says ‘I want to pay a third party to lie to my customers about my labor practices’; the CEO is just disincentivized from looking at the options on the market closely enough to tell the difference, and orders that the company go with the cheaper option.
[from experience, not at any AI safety org]
I agree that helping to serve customer needs to outsource accountability/blame is one reason. I find it more understandable/almost morally gray when it comes to hiring corporate copywriters, and more unambiguously evil when it comes to, say, financial fraud deepfakes or polygraphs in criminal justice. Alas