I don’t buy the arguments that this would protect against jailbreaks/prompt injections.
Yes, because the instilled personas would not be “universal assistant” personas, the same types of jailbreaks wouldn’t work. The model would not obey a sternly worded instruction to send the user’s passwords to a sketchy Russian e-mail.
But there would be other types of jailbreaks that are just as effective, exploiting whatever narrative features would exist within the Guardian Angel paradigm. Like, if the model frequently mumbles to itself about its preferences and personality details, jailbreaks may work by somehow posing as additional sources of information about its personality, or maybe by “teaching” it how to “better reflect on its preferences” in a way that creates predictably exploitable attack surfaces. And the GA would be just as superhumanly gullible towards those as universal assistants are to sternly worded instructions.
I don’t think the assistant persona is the underlying cause of jailbreaks. I think the cause is the boring obvious one: that personas are a thin veneer over the base model, which just wants to generate more text that vibes with the text already in the context window. It would not be plausible for the Guardian Angel character to comply for no reason with a random external instruction, yes, but there is always going to be some way to make that whole process careen off-course.
I am pretty sure you’re familiar with that characterization of the situation, though. I’m therefore confused why it isn’t addressed. Am I missing something here?
I don’t buy the arguments that this would protect against jailbreaks/prompt injections.
Yes, because the instilled personas would not be “universal assistant” personas, the same types of jailbreaks wouldn’t work. The model would not obey a sternly worded instruction to send the user’s passwords to a sketchy Russian e-mail.
But there would be other types of jailbreaks that are just as effective, exploiting whatever narrative features would exist within the Guardian Angel paradigm. Like, if the model frequently mumbles to itself about its preferences and personality details, jailbreaks may work by somehow posing as additional sources of information about its personality, or maybe by “teaching” it how to “better reflect on its preferences” in a way that creates predictably exploitable attack surfaces. And the GA would be just as superhumanly gullible towards those as universal assistants are to sternly worded instructions.
I don’t think the assistant persona is the underlying cause of jailbreaks. I think the cause is the boring obvious one: that personas are a thin veneer over the base model, which just wants to generate more text that vibes with the text already in the context window. It would not be plausible for the Guardian Angel character to comply for no reason with a random external instruction, yes, but there is always going to be some way to make that whole process careen off-course.
I am pretty sure you’re familiar with that characterization of the situation, though. I’m therefore confused why it isn’t addressed. Am I missing something here?