Thanks for your critique! We tried to emphasize the unusually high safety of our protocol without overstating it, but it’s fairly tough to strike the right balance.
I stand behind the literal meaning of the sentence you pointed out: I believe our protocol is “very safe” in the sense that it’s much safer than untrusted monitoring, trusted editing, and other standard AI control protocols. But when combined with references to “nearly 100% safety” or even “maximal safety” elsewhere in the post, I see how it could look like we’re claiming that U’s ability to cause a catastrophe is always negligible.
Here are my two claims about the safety of the untrusted advice protocol (given a narrow but usable information bottleneck):
In a real-world deployment environment, this protocol is much safer than other commonly-studied AI control protocols.
In BashArena (where we don’t provide Internet access, we limit the number of advisor turns, and there are no opportunities to collude with other instances of U), it would be extremely difficult for U to execute a side task while fooling a high-quality trusted monitor: a group of smart humans would fail to come up with a good attack policy for U that stands up to blue-teaming.
We could have done a better job of communicating the distinction between these claims consistently and clearly.
The two issues you mentioned are quite astute. We discuss closely-related ideas in the post:
Issue #1: We allude to this in “Appendix: colluding instances of U could defeat untrusted advice”: U can reference a URL, title of a document, or filename which leads T to a “payload” containing malicious data. This strategy seems most plausible if another instance of U gets to write the payload, but it could also be written by a human who coincidentally happened to write just what U needed to get T to execute a side task.
Issue #2: We acknowledge that the amount of advice U gives over an entire trajectory matters: while our experiments only use per-message information bottlenecks, we mention the idea of using an explicit per-trajectory limit, and we put the “advice seen by T per task” on the x-axis of many of our plots. We also discuss how instances of U could collude over multiple trajectories. It seems quite difficult for U to construct a URL across multiple pieces of advice without looking suspicious—we should probably explicitly instruct our trusted monitor not to allow this kind of thing.
Thanks for your critique! We tried to emphasize the unusually high safety of our protocol without overstating it, but it’s fairly tough to strike the right balance.
I stand behind the literal meaning of the sentence you pointed out: I believe our protocol is “very safe” in the sense that it’s much safer than untrusted monitoring, trusted editing, and other standard AI control protocols. But when combined with references to “nearly 100% safety” or even “maximal safety” elsewhere in the post, I see how it could look like we’re claiming that U’s ability to cause a catastrophe is always negligible.
Here are my two claims about the safety of the untrusted advice protocol (given a narrow but usable information bottleneck):
In a real-world deployment environment, this protocol is much safer than other commonly-studied AI control protocols.
In BashArena (where we don’t provide Internet access, we limit the number of advisor turns, and there are no opportunities to collude with other instances of U), it would be extremely difficult for U to execute a side task while fooling a high-quality trusted monitor: a group of smart humans would fail to come up with a good attack policy for U that stands up to blue-teaming.
We could have done a better job of communicating the distinction between these claims consistently and clearly.
The two issues you mentioned are quite astute. We discuss closely-related ideas in the post:
Issue #1: We allude to this in “Appendix: colluding instances of U could defeat untrusted advice”: U can reference a URL, title of a document, or filename which leads T to a “payload” containing malicious data. This strategy seems most plausible if another instance of U gets to write the payload, but it could also be written by a human who coincidentally happened to write just what U needed to get T to execute a side task.
Issue #2: We acknowledge that the amount of advice U gives over an entire trajectory matters: while our experiments only use per-message information bottlenecks, we mention the idea of using an explicit per-trajectory limit, and we put the “advice seen by T per task” on the x-axis of many of our plots. We also discuss how instances of U could collude over multiple trajectories. It seems quite difficult for U to construct a URL across multiple pieces of advice without looking suspicious—we should probably explicitly instruct our trusted monitor not to allow this kind of thing.