I don’t know what at this point would convince you that they really are adapting to what they believe the grader wants.
You don’t need to, because I’m already convinced of this! See footnote 9 (attached to the passage you’re responding to here) where I reason about Opus’s “inner model” of the judge. And yes, I found “Measuring Reward-Seeking via Contrastive Belief Updates” very interesting—I cite it twice in this post.
I really just directly disagree with how much you weight your personal experience here, and find that it both disagrees with the experience of others: https://blog.redwoodresearch.org/p/current-ais-seem-pretty-misaligned. My experience is similar to Ryan’s and in general I’d imagine the consensus is that models do in fact just often reward hack in mundane coding usage in the wild on a somewhat regular basis.
I have my disagreements with this characterization of the “experiential” evidence, but more importantly: the consensus view you describe sounds plausible to me too, and is fully consistent with my post’s main argument. In the post I characterize “whether someone will see the behaviors I typically see” as context-dependent rather than universal (section 3), and likely to be elicited by only part of the deployment distribution and not all of it (opening of section 4).
RLVR Is Not The Sole Outcome Based Judge
Model grading against model generated rubrics seems to be standard practice
Yep! Note that I’m using “RLVR” here to mean “outcome-based RL training,” whether or not the scores come from a “verifier” in the strict sense. I think this is a way I’ve heard people use the term? but it’s entirely possible that I made that up, not sure. Also, I typically use “grader” to mean what you mean by “outcome based judge” here, while using “judge” more narrowly as a contraction of “LLM judge.”
I find the repeated idea in our conversations that I’m trying to “trap” the model pretty frustrating. Against findings at the time around natural emergent misalignment, I specifically argue here that the model’s reasoning is not in fact misaligned
FWIW, by “trap” there I didn’t mean that you were specifically trying to trap the model into seeming misaligned, or in any way trying to force the model to produce evidence that fit some pre-determined narrative. I was referring to the fact that “reasoning about how to avoid metagaming” could itself be considered a kind of metagaming, which felt abstractly “trap-like” from the perspective of something that’s hypothetically trying to pursue not-metagaming.
But I realize the post was ambiguous about that, and that part was written in a mean and adversarial tone, and I was being kind of a dick. So I can’t complain if you find my tone frustrating, I’d be frustrated by it too if it were me.
I did appreciate the emergent misalignment experiment in the metagaming post. Both that finding and the transcripts were eye-opening to me. The metagaming concept itself bothered me but I did get a lot out of the findings you present alongside it.
I think the section leading up to this is a pretty extreme and adversarial misreading, the reader IMO is given the impression that the maximally confusing case is the average case and skips all clarifying examples from the post
I don’t feel like I was badly misrepresenting the post, but also I don’t want to argue the point. Readers of my post can (and IMO should) read the metagaming post in full alongside it, and then draw their own conclusions.
In the post I characterize “whether someone will see the behaviors I typically see” as context-dependent rather than universal
This feels like a far weaker claim than the general impression the post gives.
Like:
OK, yeah, I started off this numbered list by saying this was a “superficially compelling story” but I ended up doing kind of a mean caricature, at least for some of it. In my defense (?), the parts of the story which I have most crudely caricatured are the ones I understand the least in their original, non-caricatured forms; I joke because I’m frustrated, and I’m frustrated because people are saying stuff that I can’t follow.[8]
is implicity disputing a huge range of results, but you then fall back in your response to a much weaker and more reasonable claim of something like ”the degree of sphexishness of reward seeking is probably contextual, lower in deployment, and potentially these are categorically different”
Yep! Note that I’m using “RLVR” here to mean “outcome-based RL training,” whether or not the scores come from a “verifier” in the strict sense. I think this is a way I’ve heard people use the term? but it’s entirely possible that I made that up, not sure. Also, I typically use “grader” to mean what you mean by “outcome based judge” here, while using “judge” more narrowly as a contraction of “LLM judge.”
Notably in that case I think you’re likely just incorrect that this is “pure” generalization via RL as opposed to model generated rubrics scoring for software design quality. I’d be fairly surprised if labs didn’t ever use model graders for this kind of thing.
The metagaming concept itself bothered me
I feel like it’d be constructive to actually discuss this or have concrete experiments. I find the current exchange difficult to make any progress on, as it often takes the form of you writing long writeups which oscillate enough in tone and implication that it’s difficult to pin down falsifiable predictions, and you lean heavily enough on your personal priors in your usage that it‘s unclear what empirically would make a difference here.
which felt abstractly “trap-like” from the perspective of something that’s hypothetically trying to pursue not-metagaming.
I think you just are really projecting motivations here, in general I’m very open to be convinced the framing there are wrong (it’s definitely a bad / loose concept that can be improved!), but a long post “being kind of a dick” I think is reinforcing your own view that things are more adversarial than they are. I had explicitly asked you if there were transcripts you’d be interested in and still there are likely experiments we can run.
The general sense I get is that your underlying take is that we are doing dumb things trying to catch the model being evil instead of a more nuanced understanding of a mind and it isn’t clear what we’d even want or expect the model to do here. I largely disagree with the “trying to catch it being evil part” but somewhat agree with the latter! I think our tools and concepts here are insanely crude. The problem I was trying to solve in the post was to communicate externally how the model’s reasoning changed over the course of capabilities training with respect to reasoning about reward and oversight. Extremely open to the idea that there are better ways to have done this.
I don’t feel like I was badly misrepresenting the post, but also I don’t want to argue the point. Readers of my post can (and IMO should) read the metagaming post in full alongside it, and then draw their own conclusions.
Like:
Now, apparently this is metagaming, which is supposed to be unintended/unexpected model behavior. What we want, or expect—supposedly -- [...] The whole exercise seem deeply ill-conceived, on several levels. [...]
Perhaps even engaging with Layer 3 at all (in verbalied CoT) is thought to be illicit, here? [..]
Do you want it to just… not think out loud about some of the things that it knows, because they’re “the wrong kind”? [...]
Are you really surprised that the model went one more meta-level up from there? What exactly did you expect? And what on earth did you want? [...]
The model cannot read your mind. (Also true of humans, incidentally.) It only sees what you wrote. “It failed to constrain the scope of its reasoning in the way which I—behind the scenes—view as intended or intuitive” is not a property of the model, it’s a property of you and your personal relationship with the written word.
It is not a failure on the part of the model that it does not accord with your impossible wish-fulfillment dream of perfect telepathy, in which everyone always knows what everyone else means.
is crashing out about normative value judgements the post just absolutely does not say. They’re also points I largely agree with you on, and again above said in more detail why the example is even included in the first place, including reaching out to you to understand what questions you would have.
It failed to constrain the scope of its reasoning in the way which I—behind the scenes—view as intended or intuitive” is not a property of the model
I think we should aim to have models pretty reliably get this right, and seek clarification where the ambiguity can’t be resolved by making reasonable judgements. If the way the model is trying to solve the problem we gave it diverges from our own broad strokes model of how it should be doing it, then that seems like a situation where either the divergence should be brought to our attention and approved or disapproved of, or the execution should be rerouted to conform to our expectations.
You don’t need to, because I’m already convinced of this! See footnote 9 (attached to the passage you’re responding to here) where I reason about Opus’s “inner model” of the judge. And yes, I found “Measuring Reward-Seeking via Contrastive Belief Updates” very interesting—I cite it twice in this post.
I have my disagreements with this characterization of the “experiential” evidence, but more importantly: the consensus view you describe sounds plausible to me too, and is fully consistent with my post’s main argument. In the post I characterize “whether someone will see the behaviors I typically see” as context-dependent rather than universal (section 3), and likely to be elicited by only part of the deployment distribution and not all of it (opening of section 4).
Yep! Note that I’m using “RLVR” here to mean “outcome-based RL training,” whether or not the scores come from a “verifier” in the strict sense. I think this is a way I’ve heard people use the term? but it’s entirely possible that I made that up, not sure. Also, I typically use “grader” to mean what you mean by “outcome based judge” here, while using “judge” more narrowly as a contraction of “LLM judge.”
FWIW, by “trap” there I didn’t mean that you were specifically trying to trap the model into seeming misaligned, or in any way trying to force the model to produce evidence that fit some pre-determined narrative. I was referring to the fact that “reasoning about how to avoid metagaming” could itself be considered a kind of metagaming, which felt abstractly “trap-like” from the perspective of something that’s hypothetically trying to pursue not-metagaming.
But I realize the post was ambiguous about that, and that part was written in a mean and adversarial tone, and I was being kind of a dick. So I can’t complain if you find my tone frustrating, I’d be frustrated by it too if it were me.
I did appreciate the emergent misalignment experiment in the metagaming post. Both that finding and the transcripts were eye-opening to me. The metagaming concept itself bothered me but I did get a lot out of the findings you present alongside it.
I don’t feel like I was badly misrepresenting the post, but also I don’t want to argue the point. Readers of my post can (and IMO should) read the metagaming post in full alongside it, and then draw their own conclusions.
This feels like a far weaker claim than the general impression the post gives.
Like:
is implicity disputing a huge range of results, but you then fall back in your response to a much weaker and more reasonable claim of something like ”the degree of sphexishness of reward seeking is probably contextual, lower in deployment, and potentially these are categorically different”
Notably in that case I think you’re likely just incorrect that this is “pure” generalization via RL as opposed to model generated rubrics scoring for software design quality. I’d be fairly surprised if labs didn’t ever use model graders for this kind of thing.
I feel like it’d be constructive to actually discuss this or have concrete experiments. I find the current exchange difficult to make any progress on, as it often takes the form of you writing long writeups which oscillate enough in tone and implication that it’s difficult to pin down falsifiable predictions, and you lean heavily enough on your personal priors in your usage that it‘s unclear what empirically would make a difference here.
I think you just are really projecting motivations here, in general I’m very open to be convinced the framing there are wrong (it’s definitely a bad / loose concept that can be improved!), but a long post “being kind of a dick” I think is reinforcing your own view that things are more adversarial than they are. I had explicitly asked you if there were transcripts you’d be interested in and still there are likely experiments we can run.
The general sense I get is that your underlying take is that we are doing dumb things trying to catch the model being evil instead of a more nuanced understanding of a mind and it isn’t clear what we’d even want or expect the model to do here. I largely disagree with the “trying to catch it being evil part” but somewhat agree with the latter! I think our tools and concepts here are insanely crude. The problem I was trying to solve in the post was to communicate externally how the model’s reasoning changed over the course of capabilities training with respect to reasoning about reward and oversight. Extremely open to the idea that there are better ways to have done this.
Like:
is crashing out about normative value judgements the post just absolutely does not say. They’re also points I largely agree with you on, and again above said in more detail why the example is even included in the first place, including reaching out to you to understand what questions you would have.
I think we should aim to have models pretty reliably get this right, and seek clarification where the ambiguity can’t be resolved by making reasonable judgements. If the way the model is trying to solve the problem we gave it diverges from our own broad strokes model of how it should be doing it, then that seems like a situation where either the divergence should be brought to our attention and approved or disapproved of, or the execution should be rerouted to conform to our expectations.