However, we think it’s pretty clear that even without red-teaming, a sufficiently narrow information bottleneck makes our protocol very safe.
Strong disagree. Not only does this seem false, even if it were true it would certainly not be “pretty clear”, or it’s a motte relying on “sufficiently narrow” to mean unuseably narrow.
It seems false for multiple attack vectors which require unstated assumptions to defeat.
1) You can’t calculate how much practical information a message can refer to without knowing the context the trusted agent is in. A shortened url can point to any large amount of untrusted instructions. A smart untrusted model might have memorized many different web pages with different advice on how to code/develop, and calculate how to fool the weaker model by poisoning with “good advice” in the wrong context. Or simply give the url to a jailbreak. Information theoretically, if you don’t know what dictionary the “trusted” model has access to, you don’t know what the untrusted model is sending in its n bits.
Unstated assumptions that could fix most of this: Trusted model has no internet access, nor access to large amounts of data with different opinions
2) Real world uses are mostly multi turn, so the information accumulates, and the untrusted model can recompose instructions together, notably urls or other indexes.
Unstated assumptions that could fix most of this: Single turn only
More meta note: I liked this paper and the work itself, but the part I quote seems so weird to me and implies really bad security mindset or communication. To me it’s shooting yourself in the foot to overstate the usefulness and strength of your technique/research, as it means I can’t trust your judgement but only the arguments and results. I don’t reflectively endorse judging the author strongly based on one comment, but I wanted to communicate how bad I found that line to be.
This is fantastic coining and frame and I look forward to using it