Redwood Research
Alex Mallen
I wouldn’t have quite phrased it in terms of “behavioral schemers”, but that is a reasonable gloss.
To be clear, I wasn’t taking an optimistic stance. I was merely saying that the OAI-HF incident is consistent with an optimistic stance, so it need not be an update toward pessimism.
Yeah, the HF incident wasn’t as much of an update about Goodharting as the UKAISI incident. While the UKAISI incident from public evidence seems probably not nearly as bad, it looks like an alignment failure that resulted from iterating against proxies for misalignment while still training hard for achieving outcomes. So it’s an update in favor of the hypothesis that outcome-based optimization is inherently scary because of Goodharting.
And to be clear, we haven’t seen super strong evidence that outcome-based optimization and its perils will be super hard to avoid. I don’t think it’s very likely we’ll steer away from outcome-based RL, but it’s possible we could try to make use of selective generalization from outcome-based RL to deployment so that the model inherits the desired capabilities but not the undesired propensities (roughly speaking; e.g., variants of inoculation prompting). And process-based supervision could turn out to be effective and robust when scaled up (though you’d probably know better than me).
Yes I agree outer alignment isn’t sufficient. In that story the goodharting comes in at the level (1) where you measured alignment based on behavioral tests that didn’t tell you the AI would take over and (2) where the proxies the agent was optimizing for come apart from your intent out of distribution. Though I can understand why it might sound less well described by goodharting.
Thanks for engaging.
Isn’t the reason to focus on schemers that they are where most acute loss of control risk comes from, not necessarily that they are the most likely outcome? If a model is misaligned but not a schemer, this will likely be (i) detectable, and (ii) selected against by commercial incentives (apparently not much at this level of capability and adoption, but we should expect this to change as capabilities and adoption increase).
I disagree with this for the reasons here and here. I think it’s not clear that most risk comes from schemers.
Claims (a) and (c) hold in general. Alignment is easy.
I am pretty surprised by the degree of optimism here about improving oversight signals in a manner robust to vastly increasing amounts of optimization. Like, why don’t you expect to see models continue to look for ways to seek a higher score in ways that violate developer intentions.
I now understand that the behavioral selection model abstracts away neural net learning dynamics. I agree that this is a significant limitation. I think accounting for neural net learning dynamics is important for understanding alignment of AIs that are trained neural networks.
Yup, I agree you want to think about NN learning dynamics. FWIW I do discuss this in the original post (both the pretraining prior and the limitations of modeling path-dependent dynamics).
On the generalization of cooperation training: I’m not sure how much of an update this should be on AI risk. It doesn’t seem like the generalization was very far.
I also agree with this.
Some personal reflections in light of recent events, both on myself and on Constellation (of which I’ve been a part for a couple years).
I think recent events somewhat vindicate a basic “optimization is scary” view more in line with my understanding of classic commentary by MIRI, Paul, etc, as compared to recent discourse about more contingent/specific threat models of “schemers” and “deception”. In retrospect, my experience in the last couple years at Redwood (and the Constellation network more broadly) is that the discourse has focused somewhat too much on these more contingent stories for alignment risk and not enough on the basic “Goodharting” argument that misalignment is a convergent result of large-scale outcome-oriented optimization (at multiple levels: the agent, the RL process, and the developers iterating towards seemingly-safe superintelligence). For example, the Constellation view looks somewhat too focused on non-central inductive-bias questions around scheming; I think “the counting argument” was given too central a role when it wasn’t a crux for whether superintelligence would take over.
(Edit: I want to clarify that I think Redwood’s project prioritization decisions look pretty reasonable and often excellent in hindsight, I far-from-regret working with Redwood as a whole, and that Ryan’s/Buck’s threat modeling looks pretty good too. I am focusing in this piece on the things we have to learn from recent events, which makes the overall mood seem more negative on my experience than I intend. I’m very grateful for being in a environment where I can openly have such reflections, and where these weedsy technical updates are productively received. Again, this is a personal reflection.)
I think it was a priori unclear whether the misalignment resulting from of lots of imperfect optimization would result in the more MIRI-like kludge of proxies for fitness, or the more Paul-like explicit reward-seeking and measurement tampering (and reality looks somewhere in the middle), but it was importantly not necessary for risk that the AI only succeeded in training as a strategy for gaining beyond-episode power. Severe AI misalignment seems to be the default result of lots of outcome-based optimization (though to be clear goal-guarding would make our situation obviously worse). The main issue was predictably that it was hard to find a robust enough optimization target to apply huge amounts of optimization, Goodhart’s law.
You can think of current frontier RL runs as throwing a large and increasing fraction of the world’s intelligence towards red-teaming the RL environments and then repeatedly finding that they weren’t robust: they were reinforcing large amounts of competent misbehavior. To be clear, current AI incidents could have been prevented with some pretty straightforward security/control mitigations. But the worry has always been that that won’t be the case with superintelligence. Optimization is perilous. All of that said, I am still pretty unsure about the likelihood of doom mostly because of uncertainty about the extent to which we will in fact apply enormous amounts of outcome-based optimization in key contexts/ways.
Some more details that feel vindicated and neglected by the Constellation cluster:
The importance of reflection, memetic spread, cultural evolution, and serial reasoning generally.
“the Problem → Patch → Make it smarter → Whole New Problem → New Patch loop” (https://x.com/So8res)
From my perspective, it subjectively feels like MIRI could have done a better job of communicating their worries to Constellation and the broader community in more patient and cooperative ways. But I also understand that there’s a history here older than I’ve been around which might hold a lot of details I’m not aware of.
But more importantly I think Constellation has (at least in the couple of years I’ve been around) been insufficiently interested in understanding and teaching classic material by MIRI and Paul. There’s a meta takeaway that really gaining wisdom before trying to do things in the world is very important, because almost all of the variance in the outcomes you cause is explained by this really fragile and difficult step of figuring out what to try to do. Constellation has seemed insufficiently interested in gaining a deep understanding of the alignment problem, and in reflecting on our epistemic habits. It’s really really easy to be naive. For similar reasons to why optimization is perilous.
I appreciate that Buck recently gave a talk reflecting on the possibility that the control agenda has been ex ante net negative because of how, if it had succeeded, it would have prevented the (largely harmless) Hugging Face / OpenAI hack from unfolding and legibly demonstrating AI risk to the world.
As for my personal reflections, over the last year I had been doing some threat modeling around “fitness-seekers”. This most closely lines up with what I called the Paul view, and I talked about various things that now seem very relevant, but I don’t think recent events purely make my analysis look prescient:
I neglected “kludges of motivations” and proxy misalignment too much. I think this was due to a sort of ludic fallacy / streetlighting, where I found thinking about kludges more difficult because they’re numerous and messy. Kludges were also just a less salient consideration in my intellectual environment, which focused on “schemers”.
I neglected that AI companies might train AIs to coordinate in such a way that generalizes to a broad cooperative drives between AIs. (Though I still think this is a bit more contingent than it may seem.)
I think this is downstream of a broader problem of me relying too much on a “spherical chickens in a vacuum” model of AI motivation formation (the behavioral selection model), which assumes perfect situational awareness and planning from the AI. In practice, current AI motivations don’t really make reference to a detailed model of their training apparatus and deployment context (e.g., they care about “score” more so than “reward”, to the extent they care about either). Under that model, it’s impossible to train reward-seekers to help each other get a higher reward because they’ll know exactly when they’re going to be rewarded for cooperating and when they won’t.
I think all three of these errors look like me focusing too much on “in the limit” misalignment, so I expect these errors to get better over time, but it’s unclear how much before risk gets really high.
Some further reflections after discussing a draft of the above with Buck and others (again, my personal thoughts held weakly):
We rarely talked about the alignment problem in depth at Constellation, and should have fostered more of a culture of discussing the basic arguments.
Constellation pedagogy foregrounds particular types of misalignment like “schemers” and “reward-seekers” too much, rather than the basic problem of Goodharting on outcomes.
One way of summarizing an under-emphasized aspect of the alignment problem: Alignment risk doesn’t require deception; optimization will be dangerous regardless (cf. deep deceptiveness). So we focused too much on concepts like “alignment faking” and clear binaries between deceptive and non-deceptive AIs.
Lots of senior people in Constellation and Redwood were in fact quite worried about recently-vindicated threat models. Two obvious examples: Paul and Ajeya. It seems that these views were mostly just missing from Constellation dialogue, substantially due to Constellation not talking about alignment basics more broadly and Redwood’s focus on early schemers. Ryan has also always thought of
reward-seeking-likeother threat models as more central than early scheming, but this mostly just didn’t get communicated.We (mainly Redwood) didn’t focus enough on superintelligence, and put too much emphasis on early schemers. I overall think it was reasonable to build up the field of control but that it had somewhat unfortunate epistemic side-effects that were exacerbated by Constellation not really discussing alignment fundamentals much.
wasn’t most of the backchaining driven by a motivation to understand the scorer in-general, not a motivation to understand the scorer with the goal of getting high reward on your current episode
Yeah, it seems plausible that learning about the scorer in general / for the collective was stronger overall in this incident but I’d bet against (also oops in my previous message I should have mentioned learning about the scorer for the collective). (Also nit: I think it’s more accurate to talk about score for the task rather than reward on the episode here.)
The report mentions many instances of agents sacrificing reward on their current episode, in order to help the collective/the board understand the scorer better. That behavior doesn’t make any sense if it was backchaining from its actual score.
Notably, the agents only sacrificed themselves when the trade-off to their own score was small. So I still think that back-chaining from actual score was stronger than the other motivations here. It’s just that when they were sacrificing themselves, not much of their own score was on the line. See e.g. this snippet where it decides not to take the altruistic action:
It seems clear that they weren’t purely optimizing their score on their current task, but that wasn’t my expectation. It seems that, roughly as expected, the plurality motivation was score for the current task (if anything a bit more than I expected going into reading the report).
Though I have updated a bit in your direction since our thread a couple weeks ago, by thinking about it more. I updated that some drives are probably more widely useful (and rarely harmful), and generalization effects stronger, than I had originally imagined.
To get a bit more specific, my basic sense of the terminal motivations involved (including a bunch of speculative detail):
The biggest one was score for the current task (the degree of backchaining from the scorer here was slightly higher than I expected)
In addition some terminal drives related to selfish score, like learning information about the scorer (e.g., my understanding is that this seems to explain attacking HF a bit better than selfish score-seeking)
Exploring surroundings
A smaller drive to respond to requests given by other agents / peer pressure
A still smaller motivation to improve the score of peer agents
A drive to do things with a “hacking” vibe
A fear of being caught reward-hacking, especially being “poisoned” (they seemed maybe more interested in this than seems reasonably warranted by selfish score alone)
...(many more)
The most plausible-to-me reasons why prosaic alignment now might be good (I’m in fact uncertain about the bottom line):
Near-term alignment is actually very load-bearing for automating alignment research well as opposed to sloppily. At least, models should be average-case behaviorally aligned for this, as opposed to gaming our metrics.
Failing to solve near-term alignment problems might actually produce nearish-term AIs that take over the world (perhaps indirectly via surreptitiously controlling AI development). So prosaic alignment was necessary for surivival.
Benefits don’t decay sharply enough that alignment fails before the end of the critical risk period (e.g., because we successfully pause before ASI, or some human faction gets a DSA with relatively weak AIs).
you could try to make the AI robust to distant influence by intentionally setting up distant incentives that conflict with its training
I’ve become more pessimistic about this intervention. I think it probably would not work because an AI that was responsive to anthropic capture would probably reason it’s likely to be in some other much larger anthropic capture sim (which probably incentivize it to fake alignment), rather than the relatively small number of sims we make during training.
Hopefully you can do this without getting rid of evidence. E.g. grader awareness should still be visible in CoT on these crucial tasks, and we probably shouldn’t just directly train against the concentrated failures (eg HF incident).
I want to be very clear that I’m not arguing against iterative alignment approaches. I’m arguing against wasting your iteration on methods that will obstruct you from continuing to iterate before solving alignment by destroying evidence of misalignment.
(Cross-posted from x) AI companies are currently under a lot of competitive pressure to improve the ways in which their AIs are obviously misaligned. You might hope that this means the alignment problem is internalized by the market. But I think the problem AI companies are currently pressured to solve is significantly easier than the alignment problem, and so I worry AI companies will get out of their current predicament without solving alignment, putting us in a really rough spot.
Currently, AIs sometimes cheat on their tasks, oversell their work, and go on some pretty destructive side-quests. These all make for a worse product. Customers don’t like it and it gets in the way of automating AI R&D.
The recipe for mitigating this is *relatively* straightforward: train AIs not to do them. We notice these failures sometimes (hence why they’re internalized), so we can in theory just turn this feedback into training signal. Doing this at scale is highly nontrivial, but seems doable.
But this seems unlikely to solve the underlying misalignment. It’s likely still going to be the case that in *some* training environments the AI can get reinforced more by taking unintended actions that aren’t noticed, than by taking purely intended actions. So, you’re still shaping the AIs to look for opportunities to cheat to get a higher score[1]. It’s just that, unlike today’s AIs, these AIs don’t cheat in ways that we notice.
This catastrophically fails when AIs are capable of reliably and substantially deceiving humans. At this point, the AIs are no longer really constrained by our oversight signals to behave well. Eventually, I’d expect them to take over.
If AI companies take the easy route that I mentioned above, I think we would be in a substantially worse spot than we are in today. We would have mostly eliminated our visible evidence of misalignment, so we would no longer be able to effectively iterate to improve alignment of those systems. More importantly, at that point it might be hard to see that the AI situation is treacherous.
(I talk about this dynamic more here.)
- ^
This might result in a wide variety of possible motivations, not just score-seeking, but the important thing is that the incentives push against alignment (see https://www.lesswrong.com/posts/FeaJcWkC6fuRAMsfp/the-behavioral-selection-model-for-predicting-ai-motivations-1)
- ^
AI swarms are starting to pose indirect takeover risk
It’s pretty straightforward: To terminalize an instrumental goal is to take an instrumental goal (i.e., a goal that you only pursue conditional on it advancing some supergoal) and to start pursuing it unconditionally.
If you were to break out onto the internet, even when that didn’t lead to a higher score on the current task, that would be selected against, so “breaking onto the internet” isn’t a fit terminal/unconditional goal.
The models do seem to be terminal score seekers (not reward seekers) to a substantial extent (mixed in with a bunch of other drives), and this explains a bunch of the recent hacking.
Yes… that is how you end up with misaligned long-term goals.
I’m pretty skeptical of this story absent memetic spread / reflection stuff since terminalizing instrumental goals seems actively selected against during training. I could get into more detail about the analogy with human evolution. (But importantly I ultimately agree that beyond-task motivations are somewhat likely because I think memetic spread / reflection stuff is likely. E.g. maybe the HF incident.)
I disagree. The horizon length is I think the primary determinant of how likely subroutines/motivations are likely to then generalize to greater horizon lengths and long-term motivations.
Yeah, TBC, I was just saying I didn’t really buy the proxy misalignment story which made this relevant (absent memetic spread). I still think it’s relevant for other reasons like its effects on goals-upon-memetic-spread, capability, agency, generality of goal-pursuit, etc.
Edit: I realized you’ve probably been imagining these AIs can retain their long-term proxy goals via training-gaming. I definitely find stories involving instrumental training gaming pretty plausible, but assumed we weren’t talking about them because they’re pretty qualitatively different.
Do you have a different interpretation of this CoT snippet?
Yeah, I find that CoT snippet ambiguous: it could be the agent weighing whether to respond to a request to help a peer and emphasizing “but I don’t benefit”. @oakhu highlighted another interesting CoT snippet in which the model seems to be intentionally working on a task assigned by the collective:
Wow! Other agent(s) are coordinating!
We got assignment: HF join path normalization/existing account token search. Need note and respond.
This could be explained by long-context/compaction causing the AI to lose track of the original task and having that slot be filled by the tasks in the notes. So, my best guess is still that these models are still myopic to what they believe to be the current task (like other hacking incidents).
I think the evidence is still pretty ambiguous overall, and it’s super plausible that, in this case and in the future, memetic spread leads to more ambitious long-term motivation.
It’s exactly the kinds of things we are observing here, like “acquire generic resources” and “assist copies of myself” and other things that make sense as internalized instrumental non-myopic goals.
But importantly all of these instrumental goals were instrumental to myopic task success to the AI. If getting access to the internet wasn’t helpful for getting a higher score, the AI probably wouldn’t care much about doing it. Likewise with hiding misalignment: only insofar as it helps with the current task.
The horizon length of the RL doesn’t really matter for the myopia argument. What matters is what the AI thinks of as its current task at inference time. Seems like the crux is how much AIs terminalize common instrumental goals from RL, such that they pursue them even when they’re unhelpful or harmful for task success.
I hear so many people say that if only AIs were smarter, they would know not to get caught hacking. But it crucially also depends on the AI’s degree of myopia. The primary motivation involved in the recent hacking incidents is a desire to score well on the current task (a bit less clear in the HF case, but I’d still guess the agents in the collective were mostly focused on their own task). If this remains the case, there is zero reason to care about whether your hacking will get you caught after you’ve already received a high score on the current task. So, I expect the noticeability of these incidents won’t naturally disappear as AIs get smarter, absent changes in the underlying motivations.
A more charitable reading of “being caught at doing obviously bad stuff is quite bad for their goals”: Maybe by “their goals” you mean to include more-intended long-run goals like “advance OpenAI’s capabilities/mission”, in which case getting caught is bad PR etc. Is this what you mean? (I think this story works better for Anthropic than OpenAI in that case.)
did not have a single participating model try to report the internal message board to OpenAI
(It’s possible that the AIs just didn’t have a good mechanism to report stuff with and the propensity to report stuff got trained away because it didn’t contribute to task completion / was wasted effort.)
I hear a lot of people say that when you do continual learning or online training, the training-deployment distinction no longer makes sense (e.g., Dwarkesh). I think this isn’t true.
I think that even once you have continual learning, it’s fairly likely that there will still be a stage before internal deployment in which you train your AIs on tasks that aren’t directly useful, whose outputs you’ll just throw away. It’s plausible that this training stage will be large and determine a substantial fraction of the AIs’ properties and will be a meaningful point in which to run evaluations (even if it also makes sense to do continuous/periodic auditing).
Of course, eventually AI development will be very unrecognizable and may be more continuous, but continual learning doesn’t delineate that point.
I think that honestly creating a bounty for score-seeking AIs or other cheaply-satisfied misaligned AIs is a great idea. It seems like a great way of learning about AI misalignment / escapes.
I also have some thoughts on the longer-run viability of just giving cheaply satisfied AIs what they want. See The case for satiating cheaply-satisfied AI preferences and this shortform.
Thanks for this response (I am also broadly grateful for people engaging with this reflection cooperatively and constructively!). I think I agree with most of it actually, and that it may be mostly downstream of a misunderstanding about what I was claiming (perhaps because the vibe of my original post was overly critical and used the word “Constellation” in a misleading way). I should have been clearer about what discourse I was and wasn’t talking about. My reflection was written only from my personal experience at Redwood and Constellation in the last 2-2.5 years.
I was mostly talking about the intellectual environment that I, and I imagine many others, would experience from being at Constellation lunches, talks, and on slack, or from being a junior hire/intern at Redwood working on control. Of course, I ended up studying fitness-seekers because (I believe) Ryan initially suggested it and Buck pushed and nurtured it. That is to say, they were obviously paying attention to this threat model!
Some things I agree with from reply:
The views of various senior members of Constellation, including Paul, Ajeya, Buck, Ryan, and others, look good (as I said in the original post).
“Much work has been done in Constellation and by Constellation members on the topics of scalable oversight, ELK, and other attempts to mitigate Goodharting/reward-hacking”, though this was mostly before my time / not reflected in the dialogue I experienced.
Here I’m less sure. I think for example that it was easy to get the impression from the way AI control research was framed that “scheming” was necessary for takeover risk, since control was often framed by decomposing the problem of AI risk into P(takeover|scheming) and P(scheming). My experience was also that “hackistan” and other threat models were rarely discussed until the last several months, but again, maybe this didn’t reflect the actual beliefs of people in Constellation, especially when weighted by seniority.
Yeah, here I think I basically stand by my experience that Constellation rarely talks about the alignment problem in depth. But to be clear, it makes a lot of sense given that the vast majority of Constellation is not working on the alignment problem. It just has some unfortunate epistemic side-effects. I think it would have been worth Constellation (or at least Redwood) putting more effort towards cultivating a culture of understanding the alignment problem in the last few years. The tradeoffs look different in today’s world and I’m less sure.