Thanks for the thought! We did try different approaches to noise scheduling like you suggest. From what we tried, adding noise only once resulted in faster distillation for the same total amount of noise added/robustness gained. However, we didn’t run comprehensive experiments on it, so it’s possible a more experimentation would provide new insights.
Thanks for the thought! We did try different approaches to noise scheduling like you suggest. From what we tried, adding noise only once resulted in faster distillation for the same total amount of noise added/robustness gained. However, we didn’t run comprehensive experiments on it, so it’s possible a more experimentation would provide new insights.