It’s a good question, from my experience LLMs can sort of improve their prompt injection defense if they’re prompted explicitly to focus on finding prompt injections. However, it’s unclear to me whether this is due to better role perception or because they’re just identifying suspicious-looking text that doesn’t match their usual output (this is related to work on CoT tampering and introspection)
It’s a good question, from my experience LLMs can sort of improve their prompt injection defense if they’re prompted explicitly to focus on finding prompt injections. However, it’s unclear to me whether this is due to better role perception or because they’re just identifying suspicious-looking text that doesn’t match their usual output (this is related to work on CoT tampering and introspection)