There’s a threat model in which LLMs are constantly passing subliminal information to one another through text, both in-context and during various kinds of fine-tuning. I think it’s ~40% likely that this has safety implications already and ~70% likely that solving this is part of the critical path to making an LLM-based singularity not kill everyone, assuming such a path exists.
On the other hand, this problem is just insanely cursed. I have no idea how on earth one would even deal with this! Various kinds of subliminal learning, phantom transfer, weird poisoning, etc. have been measured over the past year and a half, and nobody has any idea how to deal with them, we’re just constantly finding even more cursed and weird ways for this kind of poison to be introduced. Sad!
The conditions for this to work predictably seem hard enough that I’m not sure if we should worry about this as an intentional attack vector? It does show how little control we actually have over AI training though.
There’s a threat model in which LLMs are constantly passing subliminal information to one another through text, both in-context and during various kinds of fine-tuning. I think it’s ~40% likely that this has safety implications already and ~70% likely that solving this is part of the critical path to making an LLM-based singularity not kill everyone, assuming such a path exists.
On the other hand, this problem is just insanely cursed. I have no idea how on earth one would even deal with this! Various kinds of subliminal learning, phantom transfer, weird poisoning, etc. have been measured over the past year and a half, and nobody has any idea how to deal with them, we’re just constantly finding even more cursed and weird ways for this kind of poison to be introduced. Sad!
The conditions for this to work predictably seem hard enough that I’m not sure if we should worry about this as an intentional attack vector? It does show how little control we actually have over AI training though.