Yeah, exactly! Although I doubt such prompts will exist without setting up the learning algorithm to induce them. For example, you could train the model to output the same logits across different recontextualization prompts on the training distribution, while outputting different logits on other distributions.
Yeah, exactly! Although I doubt such prompts will exist without setting up the learning algorithm to induce them. For example, you could train the model to output the same logits across different recontextualization prompts on the training distribution, while outputting different logits on other distributions.