That’s interesting. I wonder if it’s possible to optimize over some prompt prefixes, similar to Zhu et al. that could produce much more non-canonical generations no matter what the prompt / task itself is. I suspect it shouldn’t be the case, as non-canonicity seems to be guided by error-correction failure rather that being encoded as “a feature”, but I might be wrong
That’s interesting. I wonder if it’s possible to optimize over some prompt prefixes, similar to Zhu et al. that could produce much more non-canonical generations no matter what the prompt / task itself is. I suspect it shouldn’t be the case, as non-canonicity seems to be guided by error-correction failure rather that being encoded as “a feature”, but I might be wrong