The thing is, bot spam isn’t generally high quality yet. On many platforms the tell of name-lotta-numbers is still a viable method for identifying a bot. And part of Anthropic’s concern with the military was that AI can de-anonymize people across multiple accounts and platforms. That capability, if Dario is correct that it exists, seems in line with the ability to identify such a distillation attack. Or at least begin to. Once it begins, then RL means the AI is likely to get better at it over time. Or am I giving AI too much credit? I’m not sure.
In response to your question, I would set thresholds. At all levels below interruption of service and/or complaints from the user base, feed it poison combined with quiet pruning of the worst offenders via muting spread in the algorithm. At the level of interrupting actual users and receiving complaints, then start blocking and banning. But that’s a hard line and I assume problems will require escalation and resolution.
The thing is, bot spam isn’t generally high quality yet. On many platforms the tell of name-lotta-numbers is still a viable method for identifying a bot. And part of Anthropic’s concern with the military was that AI can de-anonymize people across multiple accounts and platforms. That capability, if Dario is correct that it exists, seems in line with the ability to identify such a distillation attack. Or at least begin to. Once it begins, then RL means the AI is likely to get better at it over time. Or am I giving AI too much credit? I’m not sure.
In response to your question, I would set thresholds. At all levels below interruption of service and/or complaints from the user base, feed it poison combined with quiet pruning of the worst offenders via muting spread in the algorithm. At the level of interrupting actual users and receiving complaints, then start blocking and banning. But that’s a hard line and I assume problems will require escalation and resolution.