In the post I characterize “whether someone will see the behaviors I typically see” as context-dependent rather than universal
This feels like a far weaker claim than the general impression the post gives.
Like:
OK, yeah, I started off this numbered list by saying this was a “superficially compelling story” but I ended up doing kind of a mean caricature, at least for some of it. In my defense (?), the parts of the story which I have most crudely caricatured are the ones I understand the least in their original, non-caricatured forms; I joke because I’m frustrated, and I’m frustrated because people are saying stuff that I can’t follow.[8]
is implicity disputing a huge range of results, but you then fall back in your response to a much weaker and more reasonable claim of something like ”the degree of sphexishness of reward seeking is probably contextual, lower in deployment, and potentially these are categorically different”
Yep! Note that I’m using “RLVR” here to mean “outcome-based RL training,” whether or not the scores come from a “verifier” in the strict sense. I think this is a way I’ve heard people use the term? but it’s entirely possible that I made that up, not sure. Also, I typically use “grader” to mean what you mean by “outcome based judge” here, while using “judge” more narrowly as a contraction of “LLM judge.”
Notably in that case I think you’re likely just incorrect that this is “pure” generalization via RL as opposed to model generated rubrics scoring for software design quality. I’d be fairly surprised if labs didn’t ever use model graders for this kind of thing.
The metagaming concept itself bothered me
I feel like it’d be constructive to actually discuss this or have concrete experiments. I find the current exchange difficult to make any progress on, as it often takes the form of you writing long writeups which oscillate enough in tone and implication that it’s difficult to pin down falsifiable predictions, and you lean heavily enough on your personal priors in your usage that it‘s unclear what empirically would make a difference here.
which felt abstractly “trap-like” from the perspective of something that’s hypothetically trying to pursue not-metagaming.
I think you just are really projecting motivations here, in general I’m very open to be convinced the framing there are wrong (it’s definitely a bad / loose concept that can be improved!), but a long post “being kind of a dick” I think is reinforcing your own view that things are more adversarial than they are. I had explicitly asked you if there were transcripts you’d be interested in and still there are likely experiments we can run.
The general sense I get is that your underlying take is that we are doing dumb things trying to catch the model being evil instead of a more nuanced understanding of a mind and it isn’t clear what we’d even want or expect the model to do here. I largely disagree with the “trying to catch it being evil part” but somewhat agree with the latter! I think our tools and concepts here are insanely crude. The problem I was trying to solve in the post was to communicate externally how the model’s reasoning changed over the course of capabilities training with respect to reasoning about reward and oversight. Extremely open to the idea that there are better ways to have done this.
I don’t feel like I was badly misrepresenting the post, but also I don’t want to argue the point. Readers of my post can (and IMO should) read the metagaming post in full alongside it, and then draw their own conclusions.
Like:
Now, apparently this is metagaming, which is supposed to be unintended/unexpected model behavior. What we want, or expect—supposedly -- [...] The whole exercise seem deeply ill-conceived, on several levels. [...]
Perhaps even engaging with Layer 3 at all (in verbalied CoT) is thought to be illicit, here? [..]
Do you want it to just… not think out loud about some of the things that it knows, because they’re “the wrong kind”? [...]
Are you really surprised that the model went one more meta-level up from there? What exactly did you expect? And what on earth did you want? [...]
The model cannot read your mind. (Also true of humans, incidentally.) It only sees what you wrote. “It failed to constrain the scope of its reasoning in the way which I—behind the scenes—view as intended or intuitive” is not a property of the model, it’s a property of you and your personal relationship with the written word.
It is not a failure on the part of the model that it does not accord with your impossible wish-fulfillment dream of perfect telepathy, in which everyone always knows what everyone else means.
is crashing out about normative value judgements the post just absolutely does not say. They’re also points I largely agree with you on, and again above said in more detail why the example is even included in the first place, including reaching out to you to understand what questions you would have.
This feels like a far weaker claim than the general impression the post gives.
Like:
is implicity disputing a huge range of results, but you then fall back in your response to a much weaker and more reasonable claim of something like ”the degree of sphexishness of reward seeking is probably contextual, lower in deployment, and potentially these are categorically different”
Notably in that case I think you’re likely just incorrect that this is “pure” generalization via RL as opposed to model generated rubrics scoring for software design quality. I’d be fairly surprised if labs didn’t ever use model graders for this kind of thing.
I feel like it’d be constructive to actually discuss this or have concrete experiments. I find the current exchange difficult to make any progress on, as it often takes the form of you writing long writeups which oscillate enough in tone and implication that it’s difficult to pin down falsifiable predictions, and you lean heavily enough on your personal priors in your usage that it‘s unclear what empirically would make a difference here.
I think you just are really projecting motivations here, in general I’m very open to be convinced the framing there are wrong (it’s definitely a bad / loose concept that can be improved!), but a long post “being kind of a dick” I think is reinforcing your own view that things are more adversarial than they are. I had explicitly asked you if there were transcripts you’d be interested in and still there are likely experiments we can run.
The general sense I get is that your underlying take is that we are doing dumb things trying to catch the model being evil instead of a more nuanced understanding of a mind and it isn’t clear what we’d even want or expect the model to do here. I largely disagree with the “trying to catch it being evil part” but somewhat agree with the latter! I think our tools and concepts here are insanely crude. The problem I was trying to solve in the post was to communicate externally how the model’s reasoning changed over the course of capabilities training with respect to reasoning about reward and oversight. Extremely open to the idea that there are better ways to have done this.
Like:
is crashing out about normative value judgements the post just absolutely does not say. They’re also points I largely agree with you on, and again above said in more detail why the example is even included in the first place, including reaching out to you to understand what questions you would have.