I think messing with the token distribution at inference time will hurt performance by pushing output (farther) out of distribution, but if the watermarking was part of training it might fine.
OK, I’m glad to hear that my opinion isn’t completely conventional! To be clear, my prediction is that it can be done with ~zero loss of performance, at inference time only. without being part of training.
I think messing with the token distribution at inference time will hurt performance by pushing output (farther) out of distribution, but if the watermarking was part of training it might fine.
OK, I’m glad to hear that my opinion isn’t completely conventional! To be clear, my prediction is that it can be done with ~zero loss of performance, at inference time only. without being part of training.