I think this is a genuine concern with BCT/ACT when we have more than 2 related inputs for the same samples (eg., multiple biases, such as in Appendix F of the paper), but not the main paper results, which only train on 1 biased setting. The result I’d expect here is mode collapse, where the LLM is trained to output a narrower distribution of responses, which RMCT largely avoids (it arguably still trains on a narrower distribution than the control since rate matching may select for more similar responses).
Thanks for the comment!
I think this is a genuine concern with BCT/ACT when we have more than 2 related inputs for the same samples (eg., multiple biases, such as in Appendix F of the paper), but not the main paper results, which only train on 1 biased setting. The result I’d expect here is mode collapse, where the LLM is trained to output a narrower distribution of responses, which RMCT largely avoids (it arguably still trains on a narrower distribution than the control since rate matching may select for more similar responses).