Real-time monitoring of the agent’s outputs would suffice with very high probability. The monitor would notice the attack and stop it; the model doesn’t have the capability (or propensity, probably) to jailbreak the monitor or obfuscate its actions.
Also maybe harmlessness/refusal training, less robustly.
What safeguards do you think would have been sufficient to prevent this?
Real-time monitoring of the agent’s outputs would suffice with very high probability. The monitor would notice the attack and stop it; the model doesn’t have the capability (or propensity, probably) to jailbreak the monitor or obfuscate its actions.
Also maybe harmlessness/refusal training, less robustly.