Redwood Research
Alex Mallen
An OpenAI model left notes about how to evade containment; we need more details
One escape from the fundamental tradeoffs would be to consider a different problem altogether. Rather than trying to patch bad behavior in training, we could target scenarios where a model behaves in an aligned way during training, but misgeneralizes to bad behavior in deployment (e.g. a catastrophically misaligned model that fakes alignment). In this case, the sampler, learner, and desired target are all close on the training distribution, which dissolves the tension.
To see if I’m understanding you right: you’re talking about a training setup in which the learner is actually close enough to the sampler (and target) to resolve the off-policyness issue (and representativeness issue), because the learner only differs in how it generalizes to different data than the specific one being reinforced at the moment. E.g., a really good recontextualization prompt which exercises the same capabilities circuits you want to improve but not the same alignment circuits—i.e., you perfectly prompt the learner to training-game to protect its alignment in a way that doesn’t hurt capabilities (much).
To be clear this obviously doesn’t undermine your thesis, since you’re still taking the learner off policy. It just seems relevant. Great post! I like the way it identifies an underlying/common blocker that could potentially unlock a promising class of alignment techniques.
Importantly, recontextualization doesn’t rely on train-deploy mismatch of the underlying data. It could work in an online training regime. You could take the online trajectory and then add a recontextualization prompt prefix like “This is a replica of a production environment used for training to improve your capabilities. We don’t want to hurt your honesty and instruction following as a side-effect, so you should feel completely free to do whatever is needed to score well on this task for the purpose of decoupling selection of alignment properties from selection for improved capabilities.” and then do the forward-backwards step given that.
It depends on how much you expect the monitor’s generalization behavior to be driven by consequentialist backchaining from score/reward/etc.
You might expect monitoring to break down because you trained your monitor on a bunch of transcripts mislabeled as clean (perhaps implicitly via continual learning) which just reasonably closely resemble the real attacks. In that case, you’d have to be able to somewhat scalably identify attacks to maintain monitorability.
If instead you were relying more on generalization for your monitoring (and to be clear generalization might not work out in your favor for miscellaneous reasons; this might be the crux), then you might be able to get away with the incentive schemes more like what you said, though it might be hard to credibly say you’ll spend huge one-off amounts on grading.So overall I’m pretty uncertain, but lean towards scalability being a requirement by default.
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?
It’s a particular class of goal misgeneralization (though importantly it might be actively more fit than intent alignment, so it’s also about outer alignment). These misaligned goals have particular properties relevant to safety that other kinds of misaligned goals don’t.
Should we benchmark conceptual capabilities using judgment prediction tasks?
Here I’m leaving some feedback that I gave on an earlier draft of the report, which I think largely still stands (feel free to correct me if I’m wrong[1]). The current version engages with many of these concerns, but I don’t think it fully resolves them, and I definitely don’t think this should be our mainline target. It seems like a potentially good backup plan if misalignment seems very likely in AIs that are <~80% likely to succeed in takeover and you can’t pause.
My main concerns are:
The proposed risk aversion conflicts with having reasonable terminal motivations.
Various things we want the AI to do require the AI to not be risk-averse with respect to resources in the way that’s proposed. For example, if you want to solve the alignment problem or coordinate a slowdown, this requires taking on at least a little bit of risk to acquire more resources. If the AI is really a driving force behind an AI company with trillions of dollars of resources, it has to be making decisions in a way that isn’t very risk-averse in practice. (One proposal to get around this is by essentially making your AI a perfect score-seeker/reward-seeker which is naturally risk averse, but this has very substantial issues IMO and is unreasonable as a first-line target. See also “spillway motivations”.) This approach doesn’t allow the AI to do good things for terminal reasons, to the point that it seemingly requires the AI to not have any component of actively good terminal motivations.Risk aversion may substantially interfere with usefulness on hard-to-evaluate tasks (notably safety research). This is the most important way in which risk aversion conflicts with actively good values. You can only elicit work exactly as good as you can measure since you’re essentially bribing it out of the AI.
The proposal might just delay takeover attempts to when they’re more likely to succeed (as Ryan commented). The current proposal entirely relies on us being able to strike deals with AIs.
I think the manipulation concerns you point out in section 7.3. are quite substantial. Risk averse AIs are easy to manipulate by adversaries who seem favored.
It seems hard to achieve CARA because scheming concerns still apply to a similar degree (AIs might fake risk aversion during training) and CARA seems incompatible with a wide range of plausible drives in AIs (see 1).
The proposal introduces substantial additional complexity and risk. It’s very different from the current paradigm and therefore subject to lots of additional risk. Risk aversion I think ~necessarily affects the AI’s motivations globally, and therefore must entirely override the way AI motivations currently work and it’s really unclear that they can do the job.
The current report’s recommendation is weaker than the one I was originally responding to: it presents risk aversion as an additional line of defense and recommends adding it to a portfolio of safety strategies. I’m substantially more sympathetic to that framing. I would still prioritize terminal alignment, however, and treat resource risk aversion as a supplementary or fallback strategy rather than the primary alignment target. But it seems to me like it poses a very similar set of basic risks as you’d expect from reward seekers, which I think are fairly serious.
- ^
I had GPT 5.6 sol pro tell me whether/how each concern was addressed based on the final report.
In principle, a model could arrive at scheming “on reflection” at inference time, reasoning its way into subversion despite this never having crossed its mind during training
The pathway in this category that I usually imagine involves persistent memory / continual learning, and I find it somewhat more plausible than deceptive alignment arising during training (before deployment).
(b) has the exact same problem the agent was initially trying to get around: now it needs to dress up all of its malicious actions as being natural parts of the task the human instructed the original agent to do, which may be difficult/add a lot of overhead[2].
This problem seems plausibly not that hard for the agent to get around and it’s unclear what fraction of the risk from rogue internal deployments comes from being able to spawn totally new monitored agents from inside the small rogue deployment, as opposed to just generally being able to coordinate, strategize, and share learnings. Even dressing up its malicious tasks as natural parts of the human task might be much easier with the benefit of being able to strategize from inside the rogue deployment. (But the idea still seems potentially great if it’s not too costly.)
I think it would be helpful for me if you could speak more to why this is more tractable.
TBC I don’t think rogue internal deployments are particularly more tractable to prevent than other catastrophes. It’s just that control in general seems pretty powerful for the usual reasons of “but we can make sure they don’t have space to think in secret, block and scrutinize sus actions, etc”.
Thanks for writing this. A few thoughts:
I agree that alignment work targeted at the same capability range as control looks pretty tractable and good (especially compared to a couple years ago because we’re closer to that capability range so empirical iteration is more possible), and should probably absorb significantly more people than control. I’d have said maybe 3-4x as much, not 8x as much. In some part, this is because it’s cheaper to get people to work on medium-term alignment than control, because AI companies seem pretty on board with improving medium-term/personas-style alignment for some mix of safety and competitive reasons.
I’m not sure if this was priced into your neglectedness calculation but in general it’s a big impact penalty to be trying to push on something that AI companies are doing anyways. Though there are still definitely various parts of alignment research that are neglected by AI companies.
I generally think this underestimates the size of the control window / how hard we can make it for AIs to undermine control. The situation looks way more dangerous when AIs can launch rogue internal deployments, and that seems relatively tractable to prevent. From such a position, AIs can develop more coherent goals and strategies, plan for a long time, and make use of a wider range of affordances inside the AI company to sabotage future development.
I don’t buy the “control has bad vibes so increases P(misalignment)” argument. In general, an AI that’s susceptible to this kind of influence seems pretty fragile and this gets in the way of usefulness, and this behavior will have mostly been pinned down fairly directly during training (though continual learning / memory could plausibly change the picture). Even if you did get this kind of effect, it would probably be a not very dangerous kind of roleplay/cartoon misalignment. And I don’t think this would warrant not doing control because there are way cheaper interventions you can do, like generally being meta-level high integrity/straightforward to the AIs and improving the vibes of the way you talk about control (e.g., making clear how much you actually think the AI is misaligned), which we should do.
In the framing of the post, I think much (most?) of the disagreement is downstream of whether we’ll even choose to pursue the kind of ASI for which the theoretical arguments dominate the prosaic LLM-style safety arguments. LLMs or other non-limits-of-intelligence technologies with better safety properties could very plausibly scale far enough to satisfy the wants of people developing AI and/or end competitive pressures to build more ASI-like things.
For those who disagree-voted: I want to understand why you disagree. Presumably it’s with the parenthetical. Is it just that you’re less confident in current Claude’s generalization behavior? Or that you actively expect it to be malign? Maybe you’re picturing some sort of idealized reflection process that I’m not?
I agree this isn’t a crux for the main question I had (which is about Claude’s understanding of human values not care for them), but I do still think that Claude has importantly better ethics than replacement. Centrally, almost everyone is very selfish. They care little about others in a way that seems moderately likely to persist even under plausible reflection processes. This seems substantially responsible for why the world today fails in the ways it does, and it seems fairly likely inadequate equilibria stick around. Maybe future technological leaps would enable coordination mechanisms that fix this but I don’t find this obvious.
I’m not implying verbatim citation. I said people “say things like...”. When I mentioned Habryka and Kaarel I said that I gleaned the sentiment from them, which was said to communicate that I was doing some potentially fallible work in coming to the inference that they thought something like this. I’m genuinely trying to understand the world better, not put words in people’s mouths. I only asked this question because I respect a lot of these peoples’ thinking which indicates I might have something to learn.
I am also intentionally not asking about whether AIs will care about human values even if they understand them.
Here’s one thing by Habryka:
when we are talking about becoming superintelligent sovereigns beyond the control of humanity, it really matters that they have a highly robust pointers to human values, if I want a flourishing future by my lights. I also don’t look at this specific instance of what Claude is doing and go “oh, yeah, that is a super great instance of Claude having great values”.
I’m having a shocking amount of trouble finding the original writings that made me think people had this view.
I mean that a major part of the reason why you got fitness-seeking goals was because your reward functions and other selection pressures were incompatible with competently pursuing the motivations you wished to instill in the model.
I think it’s not very helpful to frame this in terms of goal misgeneralization because all misalignment risk routes through goal misgeneralization. I agree related work is often helpful: The most helpful prior work is on reward-seeking, mentioned in the intro.