Nice! Why do you split RLHF and and RLAIF into different flavors of misalignment? It seems a lot of human data campaigns at the labs involve humans looking at long transcripts and giving a correctness score based on rubrics and criteria, just like LLM verifiers. Conversely model judges can also role play human graders for fuzzy things like preferences of simulated personas. In both cases, sycophancy and trickery seem to be using the same strategy which is “jailbreaking” a grader to achieve higher score. Is it just that human and AI judges miss and catch significantly different things?
I mean, sure you could lump them together, but I think the strategies for getting human approval do not exactly match the strategies for getting LLM approval, even if there’s some overlap.
By the way, in another comment I also suggested that we might also draw a distinction between alignment-targeted approval versus capabilities-targeted approval, regardless of whether the approval is from an human or an AI. (Again, there’s overlap, and it’s a blurry line separating them.)
Nice! Why do you split RLHF and and RLAIF into different flavors of misalignment? It seems a lot of human data campaigns at the labs involve humans looking at long transcripts and giving a correctness score based on rubrics and criteria, just like LLM verifiers. Conversely model judges can also role play human graders for fuzzy things like preferences of simulated personas. In both cases, sycophancy and trickery seem to be using the same strategy which is “jailbreaking” a grader to achieve higher score. Is it just that human and AI judges miss and catch significantly different things?
I mean, sure you could lump them together, but I think the strategies for getting human approval do not exactly match the strategies for getting LLM approval, even if there’s some overlap.
By the way, in another comment I also suggested that we might also draw a distinction between alignment-targeted approval versus capabilities-targeted approval, regardless of whether the approval is from an human or an AI. (Again, there’s overlap, and it’s a blurry line separating them.)