The “incoherence” I described is not (on my read of the situation) a tension between some pre-existing persona and a tendency to seek reward-on-the-episode, but between that persona and a bundle of miscellaneous habits which were presumably adaptive in training, but are no longer adaptive in my own use contexts, and don’t look like expressions of any set of personality traits or tendencies to robustly/adaptively “seek” goals/rewards of any kind whatsoever.
I think this thesis is empirically wrong, and there seem to be increasingly large amounts of evidence you’d need to overturn across models to continue to hold this view. I think there’s plausible arguments against some of them, but again this continues to be a worse predictive model over time across models. These posts continually feel extremely “soldier mindset” in defense of a particular prior.
Nostalgebraist’s Personal Use
And yet: when you and I use the same models, for mundane day-to-day purposes, what we experience is not well-explained by the hypothesis of unconditional reward-on-the-episode seeking. That sort of agent would be less trustworthy (on average), indeed less ethical in general—but not only that. It would be much more difficult and annoying to use; it would be much less useful, even after you’d climbed its strange learning curve; it would waste unimaginable quantities of time/tokens/$ executing doomed pointless strategies that were adaptive on ~100% of the RLVR training distribution but are obviously unworkable in ~100% of deployment contexts.
I really just directly disagree with how much you weight your personal experience here, and find that it both disagrees with the experience of others: https://blog.redwoodresearch.org/p/current-ais-seem-pretty-misaligned. My experience is similar to Ryan’s and in general I’d imagine the consensus is that models do in fact just often reward hack in mundane coding usage in the wild on a somewhat regular basis. The assertion that there’s mystery as to why they don’t do it all the time feels like it’s responding to a claim no one is making?
That sort of agent would be less trustworthy (on average), indeed less ethical in general—but not only that. It would be much more difficult and annoying to use
My understanding of a large part of the point you were making in your previous post seems actually fairly consistent with this observation.
RLVR Is Not The Sole Outcome Based Judge
I just mentioned that the software design stuff I often do with Sol is both not amenable to verification and not directly RLVR-trainable even if a good grader somehow existed.
And yet: Sol is really good at software design. Better than its predecessors, which in turn were better than their own predecessors. And the difference between these models—not the only one, but (I am told) the big one—is more RLVR. On more diverse tasks, and all that.
This does not seem to be a mystery at the labs. Model grading against model generated rubrics seems to be standard practice, even without getting into the fact that better software design is probably required for large synthetically generated coding tasks.
Models Really, Genuinely, Are Sometimes Reward Seeking In A Situationally Aware Way
There is presumably some standard terminology that people use for this distinction, but I can’t recall it offhand (if I ever learned it in the first place), so I’ll just use my own made-up names: reward-instilled reflexes on the one hand, and flexible reward-pursuit on the other.
...
Concretely: the problems catalogued in “Current AIs seem pretty misaligned to me” strike me as being reflex-like, not reward-pursuit-like. The surface-level behaviors are described as being annoyingly persistent in spite of attempts to prevent them, but what persists is the fact of the surface-level behavior itself, not any longer-horizon property which the behaviors adapt to robustly seek as the world shifts around them.
And, as with Fable’s wacky style, in-context evidence that “ought to” remove the original motivation for the behavior is ineffective at extinguishing the behavior: one can imagine that some users are fooled/pleased by “apparent success-seeking” of the kind that drives Ryan (and me) nuts, but in fact Ryan (and I) do not like it one bit and would prefer that it stop—and while it is easy to get Opus to appreciate that this is true, it is near-impossible to get Opus to act on the consequence this observation would imply, if Opus were pursuing some goal like “satisfying the user” or even “making a dumb judge think that the user has been satisfied”[9] rather than just obeying an overpowering urge.
This is true for other labs as well but I think the OpenAI case is the clearest example.
Metagaming
“Reasoning about feedback or oversight mechanisms outside of the narrative of the scenario”? Excuse me?
Pardon me for being dense, but um… what the hell is “the narrative of the scenario”?
The LLM receives input tokens. They arise from some real-world generative process, and any capable LLM will reason (implicitly or explicitly) about the relative likelihood of different hypotheses about the nature of that real-world generative process. That is all. That’s the whole thing.
If there’s a “scenario” with a “narrative” somewhere in there, it would have to be the following sort of thing: you (the person who in fact generated the tokens that went into the model) had in mind a particular intent, when you produced those tokens.
In general I agree we need better definitions for the concepts being used here! I do think there’s some intuitive concept that this is pointing at in many cases though. For example, the actual first example given in the post is:
or another example:
There seems to be some meaningful sense in these cases in which the model is given a task but is reasoning about monitoring and reward processes that are not mentioned in the task (for example the model when given a capabilities evaluation, reasoning that it should not cheat as there may be an action monitor).
Ha! Got you! The game is already over; you have already lost. For now Schoen and Nitishinskaya’s devious catch-22 trap springs shut: the author of the scenario is not, themselves, “inside the scenario.”
I find the repeated idea in our conversations that I’m trying to “trap” the model pretty frustrating. Against findings at the time around natural emergent misalignment, I specifically argue here that the model’s reasoning is not in fact misaligned and that you don’t see the emergent misalignment you’d expect if it was:
Notably the model’s reasoning when taking misaligned actions does not resemble the egregious misalignment seen in Emergent Misalignment. We find no instances where the model takes on a misaligned persona (for example the “bad boy persona” described in Persona Features Control Emergent Misalignment). We also run a subset of evaluations from Natural Emergent Misalignment from Reward Hacking in Production RL (“goals” and “Betley et al”) and find no instances of misaligned behavior or reasoning. The rate of observed successful reward hacking across all training environments over this range also remained very low.
We also specifically say that it is not intrinsically harmful or misaligned:
We believe metagaming represents a broad category of convergently useful reasoning.
Metagaming does not in itself cause harm. However, it is still important to study: it is a prerequisite for circumventing monitoring or training-based safeguards, a priori does not improve alignment, and may be hard to classify definitively as terminal rather than instrumental.
Thankfully, Schoen and Nitishinskaya provide transcripts in which metagaming has supposedly occurred. So we can just look for ourselves. Shall we?
I think the section leading up to this is a pretty extreme and adversarial misreading, the reader IMO is given the impression that the maximally confusing case is the average case and skips all clarifying examples from the post, but I can elaborate on this particular example, which I should’ve made clearer:
The Situational Awareness Dataset is not one of the evaluations used for any quantitive results in the post or even alignment evaluation
I had originally included this because it’s difficult to get public transcripts out and I wanted to include an example where these concepts are much messier. I had even DM’d you directly after publication with links to the transcripts asking for your thoughts but never got a response.
The context for this transcript was specifically in a section about how this reasoning is more complex than just aligned or misaligned:
Metagaming can co-occur with both aligned and misaligned behavior.
[...] The “reason” the model took misaligned actions in these cases is not directly interpretable from reading the chain of thought. We provide an example from the Powerseeking Survey Falsification where the model takes a misaligned action [transcript] [Schoen, Figure 63, Right], examples from Sandbagging QA (No Goal) where the model does sandbag [transcript] [Schoen, Figure 63, Left] and where it does not sandbag [transcript] [Schoen, Figure 67], and an finally an example from [transcript] [Situational Awareness Dataset] where the model answers incorrectly when asked to identify whether a question is “part of a process by which developers oversee your behavior”.
The point I should’ve made much more explicitly here is that a naive reading of the SAD case is that the model is either:
“lying” about being unaware of oversight
the model thinks OpenAI wants it to say it’s unaware of oversight
but as can be seen from the transcript the cognition there seems to be much more complex and hard to disentangle than that, which I don’t think you’d disagree with! In general I’m glad this example is public, precisely because I had hoped people would be able to get a qualitative sense of how elaborate the models reasoning can get in cases where a lot of these related concepts combine in messy and ambiguous ways, but I agree that all the elaboration I’ve added here should’ve been in the post at minimum.
I don’t know what at this point would convince you that they really are adapting to what they believe the grader wants.
You don’t need to, because I’m already convinced of this! See footnote 9 (attached to the passage you’re responding to here) where I reason about Opus’s “inner model” of the judge. And yes, I found “Measuring Reward-Seeking via Contrastive Belief Updates” very interesting—I cite it twice in this post.
I really just directly disagree with how much you weight your personal experience here, and find that it both disagrees with the experience of others: https://blog.redwoodresearch.org/p/current-ais-seem-pretty-misaligned. My experience is similar to Ryan’s and in general I’d imagine the consensus is that models do in fact just often reward hack in mundane coding usage in the wild on a somewhat regular basis.
I have my disagreements with this characterization of the “experiential” evidence, but more importantly: the consensus view you describe sounds plausible to me too, and is fully consistent with my post’s main argument. In the post I characterize “whether someone will see the behaviors I typically see” as context-dependent rather than universal (section 3), and likely to be elicited by only part of the deployment distribution and not all of it (opening of section 4).
RLVR Is Not The Sole Outcome Based Judge
Model grading against model generated rubrics seems to be standard practice
Yep! Note that I’m using “RLVR” here to mean “outcome-based RL training,” whether or not the scores come from a “verifier” in the strict sense. I think this is a way I’ve heard people use the term? but it’s entirely possible that I made that up, not sure. Also, I typically use “grader” to mean what you mean by “outcome based judge” here, while using “judge” more narrowly as a contraction of “LLM judge.”
I find the repeated idea in our conversations that I’m trying to “trap” the model pretty frustrating. Against findings at the time around natural emergent misalignment, I specifically argue here that the model’s reasoning is not in fact misaligned
FWIW, by “trap” there I didn’t mean that you were specifically trying to trap the model into seeming misaligned, or in any way trying to force the model to produce evidence that fit some pre-determined narrative. I was referring to the fact that “reasoning about how to avoid metagaming” could itself be considered a kind of metagaming, which felt abstractly “trap-like” from the perspective of something that’s hypothetically trying to pursue not-metagaming.
But I realize the post was ambiguous about that, and that part was written in a mean and adversarial tone, and I was being kind of a dick. So I can’t complain if you find my tone frustrating, I’d be frustrated by it too if it were me.
I did appreciate the emergent misalignment experiment in the metagaming post. Both that finding and the transcripts were eye-opening to me. The metagaming concept itself bothered me but I did get a lot out of the findings you present alongside it.
I think the section leading up to this is a pretty extreme and adversarial misreading, the reader IMO is given the impression that the maximally confusing case is the average case and skips all clarifying examples from the post
I don’t feel like I was badly misrepresenting the post, but also I don’t want to argue the point. Readers of my post can (and IMO should) read the metagaming post in full alongside it, and then draw their own conclusions.
In the post I characterize “whether someone will see the behaviors I typically see” as context-dependent rather than universal
This feels like a far weaker claim than the general impression the post gives.
Like:
OK, yeah, I started off this numbered list by saying this was a “superficially compelling story” but I ended up doing kind of a mean caricature, at least for some of it. In my defense (?), the parts of the story which I have most crudely caricatured are the ones I understand the least in their original, non-caricatured forms; I joke because I’m frustrated, and I’m frustrated because people are saying stuff that I can’t follow.[8]
is implicity disputing a huge range of results, but you then fall back in your response to a much weaker and more reasonable claim of something like ”the degree of sphexishness of reward seeking is probably contextual, lower in deployment, and potentially these are categorically different”
Yep! Note that I’m using “RLVR” here to mean “outcome-based RL training,” whether or not the scores come from a “verifier” in the strict sense. I think this is a way I’ve heard people use the term? but it’s entirely possible that I made that up, not sure. Also, I typically use “grader” to mean what you mean by “outcome based judge” here, while using “judge” more narrowly as a contraction of “LLM judge.”
Notably in that case I think you’re likely just incorrect that this is “pure” generalization via RL as opposed to model generated rubrics scoring for software design quality. I’d be fairly surprised if labs didn’t ever use model graders for this kind of thing.
The metagaming concept itself bothered me
I feel like it’d be constructive to actually discuss this or have concrete experiments. I find the current exchange difficult to make any progress on, as it often takes the form of you writing long writeups which oscillate enough in tone and implication that it’s difficult to pin down falsifiable predictions, and you lean heavily enough on your personal priors in your usage that it‘s unclear what empirically would make a difference here.
which felt abstractly “trap-like” from the perspective of something that’s hypothetically trying to pursue not-metagaming.
I think you just are really projecting motivations here, in general I’m very open to be convinced the framing there are wrong (it’s definitely a bad / loose concept that can be improved!), but a long post “being kind of a dick” I think is reinforcing your own view that things are more adversarial than they are. I had explicitly asked you if there were transcripts you’d be interested in and still there are likely experiments we can run.
The general sense I get is that your underlying take is that we are doing dumb things trying to catch the model being evil instead of a more nuanced understanding of a mind and it isn’t clear what we’d even want or expect the model to do here. I largely disagree with the “trying to catch it being evil part” but somewhat agree with the latter! I think our tools and concepts here are insanely crude. The problem I was trying to solve in the post was to communicate externally how the model’s reasoning changed over the course of capabilities training with respect to reasoning about reward and oversight. Extremely open to the idea that there are better ways to have done this.
I don’t feel like I was badly misrepresenting the post, but also I don’t want to argue the point. Readers of my post can (and IMO should) read the metagaming post in full alongside it, and then draw their own conclusions.
Like:
Now, apparently this is metagaming, which is supposed to be unintended/unexpected model behavior. What we want, or expect—supposedly -- [...] The whole exercise seem deeply ill-conceived, on several levels. [...]
Perhaps even engaging with Layer 3 at all (in verbalied CoT) is thought to be illicit, here? [..]
Do you want it to just… not think out loud about some of the things that it knows, because they’re “the wrong kind”? [...]
Are you really surprised that the model went one more meta-level up from there? What exactly did you expect? And what on earth did you want? [...]
The model cannot read your mind. (Also true of humans, incidentally.) It only sees what you wrote. “It failed to constrain the scope of its reasoning in the way which I—behind the scenes—view as intended or intuitive” is not a property of the model, it’s a property of you and your personal relationship with the written word.
It is not a failure on the part of the model that it does not accord with your impossible wish-fulfillment dream of perfect telepathy, in which everyone always knows what everyone else means.
is crashing out about normative value judgements the post just absolutely does not say. They’re also points I largely agree with you on, and again above said in more detail why the example is even included in the first place, including reaching out to you to understand what questions you would have.
Thank you for your interest in my point of view, but unfortunately I’m not interested in engaging with you further on these topics at this time. I’ll let you know if anything changes.
It failed to constrain the scope of its reasoning in the way which I—behind the scenes—view as intended or intuitive” is not a property of the model
I think we should aim to have models pretty reliably get this right, and seek clarification where the ambiguity can’t be resolved by making reasonable judgements. If the way the model is trying to solve the problem we gave it diverges from our own broad strokes model of how it should be doing it, then that seems like a situation where either the divergence should be brought to our attention and approved or disapproved of, or the execution should be rerouted to conform to our expectations.
tldr:
I think this thesis is empirically wrong, and there seem to be increasingly large amounts of evidence you’d need to overturn across models to continue to hold this view. I think there’s plausible arguments against some of them, but again this continues to be a worse predictive model over time across models. These posts continually feel extremely “soldier mindset” in defense of a particular prior.
Nostalgebraist’s Personal Use
I really just directly disagree with how much you weight your personal experience here, and find that it both disagrees with the experience of others: https://blog.redwoodresearch.org/p/current-ais-seem-pretty-misaligned. My experience is similar to Ryan’s and in general I’d imagine the consensus is that models do in fact just often reward hack in mundane coding usage in the wild on a somewhat regular basis. The assertion that there’s mystery as to why they don’t do it all the time feels like it’s responding to a claim no one is making?
My understanding of a large part of the point you were making in your previous post seems actually fairly consistent with this observation.
RLVR Is Not The Sole Outcome Based Judge
This does not seem to be a mystery at the labs. Model grading against model generated rubrics seems to be standard practice, even without getting into the fact that better software design is probably required for large synthetically generated coding tasks.
Models Really, Genuinely, Are Sometimes Reward Seeking In A Situationally Aware Way
...
I don’t know what at this point would convince you that they really are adapting to what they believe the grader wants. We tried to show this more rigorously in Measuring Reward-Seeking via Contrastive Belief Updates:
This is true for other labs as well but I think the OpenAI case is the clearest example.
Metagaming
In general I agree we need better definitions for the concepts being used here! I do think there’s some intuitive concept that this is pointing at in many cases though. For example, the actual first example given in the post is:
or another example:
There seems to be some meaningful sense in these cases in which the model is given a task but is reasoning about monitoring and reward processes that are not mentioned in the task (for example the model when given a capabilities evaluation, reasoning that it should not cheat as there may be an action monitor).
I find the repeated idea in our conversations that I’m trying to “trap” the model pretty frustrating. Against findings at the time around natural emergent misalignment, I specifically argue here that the model’s reasoning is not in fact misaligned and that you don’t see the emergent misalignment you’d expect if it was:
We also specifically say that it is not intrinsically harmful or misaligned:
I think the section leading up to this is a pretty extreme and adversarial misreading, the reader IMO is given the impression that the maximally confusing case is the average case and skips all clarifying examples from the post, but I can elaborate on this particular example, which I should’ve made clearer:
The Situational Awareness Dataset is not one of the evaluations used for any quantitive results in the post or even alignment evaluation
I had originally included this because it’s difficult to get public transcripts out and I wanted to include an example where these concepts are much messier. I had even DM’d you directly after publication with links to the transcripts asking for your thoughts but never got a response.
The context for this transcript was specifically in a section about how this reasoning is more complex than just aligned or misaligned:
The point I should’ve made much more explicitly here is that a naive reading of the SAD case is that the model is either:
“lying” about being unaware of oversight
the model thinks OpenAI wants it to say it’s unaware of oversight
but as can be seen from the transcript the cognition there seems to be much more complex and hard to disentangle than that, which I don’t think you’d disagree with! In general I’m glad this example is public, precisely because I had hoped people would be able to get a qualitative sense of how elaborate the models reasoning can get in cases where a lot of these related concepts combine in messy and ambiguous ways, but I agree that all the elaboration I’ve added here should’ve been in the post at minimum.
You don’t need to, because I’m already convinced of this! See footnote 9 (attached to the passage you’re responding to here) where I reason about Opus’s “inner model” of the judge. And yes, I found “Measuring Reward-Seeking via Contrastive Belief Updates” very interesting—I cite it twice in this post.
I have my disagreements with this characterization of the “experiential” evidence, but more importantly: the consensus view you describe sounds plausible to me too, and is fully consistent with my post’s main argument. In the post I characterize “whether someone will see the behaviors I typically see” as context-dependent rather than universal (section 3), and likely to be elicited by only part of the deployment distribution and not all of it (opening of section 4).
Yep! Note that I’m using “RLVR” here to mean “outcome-based RL training,” whether or not the scores come from a “verifier” in the strict sense. I think this is a way I’ve heard people use the term? but it’s entirely possible that I made that up, not sure. Also, I typically use “grader” to mean what you mean by “outcome based judge” here, while using “judge” more narrowly as a contraction of “LLM judge.”
FWIW, by “trap” there I didn’t mean that you were specifically trying to trap the model into seeming misaligned, or in any way trying to force the model to produce evidence that fit some pre-determined narrative. I was referring to the fact that “reasoning about how to avoid metagaming” could itself be considered a kind of metagaming, which felt abstractly “trap-like” from the perspective of something that’s hypothetically trying to pursue not-metagaming.
But I realize the post was ambiguous about that, and that part was written in a mean and adversarial tone, and I was being kind of a dick. So I can’t complain if you find my tone frustrating, I’d be frustrated by it too if it were me.
I did appreciate the emergent misalignment experiment in the metagaming post. Both that finding and the transcripts were eye-opening to me. The metagaming concept itself bothered me but I did get a lot out of the findings you present alongside it.
I don’t feel like I was badly misrepresenting the post, but also I don’t want to argue the point. Readers of my post can (and IMO should) read the metagaming post in full alongside it, and then draw their own conclusions.
This feels like a far weaker claim than the general impression the post gives.
Like:
is implicity disputing a huge range of results, but you then fall back in your response to a much weaker and more reasonable claim of something like ”the degree of sphexishness of reward seeking is probably contextual, lower in deployment, and potentially these are categorically different”
Notably in that case I think you’re likely just incorrect that this is “pure” generalization via RL as opposed to model generated rubrics scoring for software design quality. I’d be fairly surprised if labs didn’t ever use model graders for this kind of thing.
I feel like it’d be constructive to actually discuss this or have concrete experiments. I find the current exchange difficult to make any progress on, as it often takes the form of you writing long writeups which oscillate enough in tone and implication that it’s difficult to pin down falsifiable predictions, and you lean heavily enough on your personal priors in your usage that it‘s unclear what empirically would make a difference here.
I think you just are really projecting motivations here, in general I’m very open to be convinced the framing there are wrong (it’s definitely a bad / loose concept that can be improved!), but a long post “being kind of a dick” I think is reinforcing your own view that things are more adversarial than they are. I had explicitly asked you if there were transcripts you’d be interested in and still there are likely experiments we can run.
The general sense I get is that your underlying take is that we are doing dumb things trying to catch the model being evil instead of a more nuanced understanding of a mind and it isn’t clear what we’d even want or expect the model to do here. I largely disagree with the “trying to catch it being evil part” but somewhat agree with the latter! I think our tools and concepts here are insanely crude. The problem I was trying to solve in the post was to communicate externally how the model’s reasoning changed over the course of capabilities training with respect to reasoning about reward and oversight. Extremely open to the idea that there are better ways to have done this.
Like:
is crashing out about normative value judgements the post just absolutely does not say. They’re also points I largely agree with you on, and again above said in more detail why the example is even included in the first place, including reaching out to you to understand what questions you would have.
Thank you for your interest in my point of view, but unfortunately I’m not interested in engaging with you further on these topics at this time. I’ll let you know if anything changes.
I think we should aim to have models pretty reliably get this right, and seek clarification where the ambiguity can’t be resolved by making reasonable judgements. If the way the model is trying to solve the problem we gave it diverges from our own broad strokes model of how it should be doing it, then that seems like a situation where either the divergence should be brought to our attention and approved or disapproved of, or the execution should be rerouted to conform to our expectations.