Models are very sophisticated now, if a model sees that this plugin is installed it can just evade or remove it. I think general purpose monitors such as Claude Code’s Auto Mode are a better solution here.
No, that’s confused. The point with sanitizing prompt injections is to protect your model’s alignment (and your monitor’s alignment) from degradation and attack by maliciously optimized text from another model. Why would an AI that starts off doing your task want to remove the plugin?
Models are very sophisticated now, if a model sees that this plugin is installed it can just evade or remove it. I think general purpose monitors such as Claude Code’s Auto Mode are a better solution here.
No, that’s confused. The point with sanitizing prompt injections is to protect your model’s alignment (and your monitor’s alignment) from degradation and attack by maliciously optimized text from another model. Why would an AI that starts off doing your task want to remove the plugin?
Also, we can do both. There’s no real tradeoff.
ETA: rephrased, focused on key counterpoints.