Strictly speaking, it’s a type of RLAF—RL from AI Feedback.
Yes, it’s used by all major labs, and it’s known to cause all kinds of degeneracy.
A lot of “guessing the teacher’s password” can get baked into the model—and with the “teacher” being a static AI target, the “student” AI can home in onto the teacher’s weaknesses and hammer onto them relentlessly. Mitigating that is a major challenge for all RLAF approaches.
Strictly speaking, it’s a type of RLAF—RL from AI Feedback.
Yes, it’s used by all major labs, and it’s known to cause all kinds of degeneracy.
A lot of “guessing the teacher’s password” can get baked into the model—and with the “teacher” being a static AI target, the “student” AI can home in onto the teacher’s weaknesses and hammer onto them relentlessly. Mitigating that is a major challenge for all RLAF approaches.