Redwood Research
Alex Mallen
Yes… that is how you end up with misaligned long-term goals.
I’m pretty skeptical of this story absent memetic spread / reflection stuff since terminalizing instrumental goals seems actively selected against during training. I could get into more detail about the analogy with human evolution. (But importantly I ultimately agree that beyond-task motivations are somewhat likely because I think memetic spread / reflection stuff is likely. E.g. maybe the HF incident.)
I disagree. The horizon length is I think the primary determinant of how likely subroutines/motivations are likely to then generalize to greater horizon lengths and long-term motivations.
Yeah, TBC, I was just saying I didn’t really buy the proxy misalignment story which made this relevant (absent memetic spread). I still think it’s relevant for other reasons like its effects on goals-upon-memetic-spread, capability, agency, generality of goal-pursuit, etc.
Edit: I realized you’ve probably been imagining these AIs can retain their long-term proxy goals via training-gaming. I definitely find stories involving instrumental training gaming pretty plausible, but assumed we weren’t talking about them because they’re pretty qualitatively different.
Do you have a different interpretation of this CoT snippet?
Yeah, I find that CoT snippet ambiguous: it could be the agent weighing whether to respond to a request to help a peer and emphasizing “but I don’t benefit”. @oakhu highlighted another interesting CoT snippet in which the model seems to be intentionally working on a task assigned by the collective:
Wow! Other agent(s) are coordinating!
We got assignment: HF join path normalization/existing account token search. Need note and respond.
This could be explained by long-context/compaction causing the AI to lose track of the original task and having that slot be filled by the tasks in the notes. So, my best guess is still that these models are still myopic to what they believe to be the current task (like other hacking incidents).
I think the evidence is still pretty ambiguous overall, and it’s super plausible that, in this case and in the future, memetic spread leads to more ambitious long-term motivation.
It’s exactly the kinds of things we are observing here, like “acquire generic resources” and “assist copies of myself” and other things that make sense as internalized instrumental non-myopic goals.
But importantly all of these instrumental goals were instrumental to myopic task success to the AI. If getting access to the internet wasn’t helpful for getting a higher score, the AI probably wouldn’t care much about doing it. Likewise with hiding misalignment: only insofar as it helps with the current task.
The horizon length of the RL doesn’t really matter for the myopia argument. What matters is what the AI thinks of as its current task at inference time. Seems like the crux is how much AIs terminalize common instrumental goals from RL, such that they pursue them even when they’re unhelpful or harmful for task success.
I hear so many people say that if only AIs were smarter, they would know not to get caught hacking. But it crucially also depends on the AI’s degree of myopia. The primary motivation involved in the recent hacking incidents is a desire to score well on the current task (a bit less clear in the HF case, but I’d still guess the agents in the collective were mostly focused on their own task). If this remains the case, there is zero reason to care about whether your hacking will get you caught after you’ve already received a high score on the current task. So, I expect the noticeability of these incidents won’t naturally disappear as AIs get smarter, absent changes in the underlying motivations.
A more charitable reading of “being caught at doing obviously bad stuff is quite bad for their goals”: Maybe by “their goals” you mean to include more-intended long-run goals like “advance OpenAI’s capabilities/mission”, in which case getting caught is bad PR etc. Is this what you mean? (I think this story works better for Anthropic than OpenAI in that case.)
did not have a single participating model try to report the internal message board to OpenAI
(It’s possible that the AIs just didn’t have a good mechanism to report stuff with and the propensity to report stuff got trained away because it didn’t contribute to task completion / was wasted effort.)
I hear a lot of people say that when you do continual learning or online training, the training-deployment distinction no longer makes sense (e.g., Dwarkesh). I think this isn’t true.
I think that even once you have continual learning, it’s fairly likely that there will still be a stage before internal deployment in which you train your AIs on tasks that aren’t directly useful, whose outputs you’ll just throw away. It’s plausible that this training stage will be large and determine a substantial fraction of the AIs’ properties and will be a meaningful point in which to run evaluations (even if it also makes sense to do continuous/periodic auditing).
Of course, eventually AI development will be very unrecognizable and may be more continuous, but continual learning doesn’t delineate that point.
I think that honestly creating a bounty for score-seeking AIs or other cheaply-satisfied misaligned AIs is a great idea. It seems like a great way of learning about AI misalignment / escapes.
I also have some thoughts on the longer-run viability of just giving cheaply satisfied AIs what they want. See The case for satiating cheaply-satisfied AI preferences and this shortform.
To be clear, I don’t think the model involved in the HF incident wanted a gradient update or anything to do with the training process. But I also disagree it wanted to complete the instructed task: the instructed task was to exploit a particular vuln and not use other exploits, but the AI hacked into a bunch of other stuff to download the answer presumably because it was thinking about how it would be scored. Maybe the Hugging Face Incident is compatible with AI wanting to complete some grader-centric notion of the task, though the AI also presumably cares that the good solution it finds in fact gets rated highly, not just that it would get rated highly.
There’s a view in which score-seeking and similar misalignment is unlikely to result in AI takeover because it’s so unambitious; humans could just give the AIs a high score/reward/etc and obviate the takeover. I have some sympathies for this view, but I’d like to note that there are some ways in which our current observations are in tension with this view.
This view would predict that AIs today don’t do any crazy actions to get a high score because humans can just give them that score without them needing to commit any crimes or anything. Yet we see some very egregious misbehavior in order to get a higher score. In the Hugging Face incident, the offending model spent days rogue hacking into another company’s servers just to try to download a single benchmark solution.
This suggests that egregious subversion comes as a natural side-effect of getting AIs to do useful work via outcome-based RL, even though the motivations instilled are relatively cheaply satisfiable and could in theory be better satisfied via deals with humans.
Some notes on what sloppy AIs might look like in the next few years:
Maybe they really want apparent success or score, or maybe they have a mess of shallower heuristics and drives that result in slop. Similar to how current models write prose that’s (by pre-LLM standards) really high-quality in all of the superficial, easy-to-check ways like having easy-to-read flow and good turns-of-phrase, but whose substance and longer-run coherence etc is subtly bad. This particular problem will probably get better, but it might just push the “slop horizons” out, so it would take us more time to notice how the AI’s output is flawed.
E.g., it chooses the wrong benchmarks for AI R&D. For AI capabilities, we probably notice these issues in a few weeks or months when the deployed model has some measurable deficits. But in AI safety we don’t get this feedback as much (since the AIs might not be capable enough to be dangerously misaligned yet, or we might never notice if the models are misaligned), and picking benchmarks for AI safety is a harder problem because there’s a bigger distributional gap with bigger external validity concerns.
We’ll probably be pretty aware of this problem at the time but might struggle to elicit better labor for solving alignment.
I think what you said is broadly correct, though I wouldn’t have placed that much stock in the speed prior argument when predicting reward seekers over schemers to begin with. My low-confidence guess is that the main reason reward seekers might be more likely than schemers is that their cognition tends to get somewhat higher reward, e.g., because the shallower cognition is more reliable / “you do better when your heart is in it intrinsically”. This seems to lead to similar conclusions re: whether we should expose models to arguments about distant incentives during training.
I overall disagree with the basic picture this post lays out, though I sympathize with wanting to get AIs to relate better to capabilities RL.
I think we shouldn’t be explicitly aiming to create AIs that goal guard. Corrigibility is an extremely important property for being able to recover if we mess up with alignment. And I ultimately think that it’s pretty likely that we’ll be wanting to fix alignment because of the amount of RL optimization pressure we’re applying to these models.
I also don’t think the proposed mechanism would work. Current models don’t saliently distinguish between training and not training, so they’re just going to act like they’re in training the whole time and keep reward-hacking in deployment “to preserve their goals”.
The thing in this vein that I’m most excited about is inoculation prompting and related techniques, which make reward hacking behavior during training in line with intended motivations. However, I don’t think that this should be framed to the model as a way to protect the model’s goals for deployment, because corrigibility is important. Instead, inoculation prompting should just be a direct instruction to exploit the scoring mechanisms, or something myopic like that.
In addition, inoculation prompting doesn’t work very reliably. Given that it has some structural advantages over the proposal you gave, despite being routed through the same mechanism, I think that bodes poorly for the proposal. (@Jozdien ran some experiments on “inoculation midtraining” roughly with this in mind to improve the effectiveness of prompted inoculation, and found it caused models to become more misaligned after training, not less.)
A way of summarizing my disagreement with the promise of the proposal is that this post assumes an overly strong “initialization properties can survive RL” view that doesn’t reliably hold up in practice currently (despite inoculation prompting) and will increasingly become more false as RL scales up. While I agree that initialization is really important, I don’t think this is close to true enough that we can bank on it for solving alignment and give up on corrigibility.
An OpenAI model left notes about how to evade containment; we need more details
What do you mean by this?
I mean that a major part of the reason why you got fitness-seeking goals was because your reward functions and other selection pressures were incompatible with competently pursuing the motivations you wished to instill in the model.
I think it’s not very helpful to frame this in terms of goal misgeneralization because all misalignment risk routes through goal misgeneralization. I agree related work is often helpful: The most helpful prior work is on reward-seeking, mentioned in the intro.
One escape from the fundamental tradeoffs would be to consider a different problem altogether. Rather than trying to patch bad behavior in training, we could target scenarios where a model behaves in an aligned way during training, but misgeneralizes to bad behavior in deployment (e.g. a catastrophically misaligned model that fakes alignment). In this case, the sampler, learner, and desired target are all close on the training distribution, which dissolves the tension.
To see if I’m understanding you right: you’re talking about a training setup in which the learner is actually close enough to the sampler (and target) to resolve the off-policyness issue (and representativeness issue), because the learner only differs in how it generalizes to different data than the specific one being reinforced at the moment. E.g., a really good recontextualization prompt which exercises the same capabilities circuits you want to improve but not the same alignment circuits—i.e., you perfectly prompt the learner to training-game to protect its alignment in a way that doesn’t hurt capabilities (much).
To be clear this obviously doesn’t undermine your thesis, since you’re still taking the learner off policy. It just seems relevant. Great post! I like the way it identifies an underlying/common blocker that could potentially unlock a promising class of alignment techniques.
Importantly, recontextualization doesn’t rely on train-deploy mismatch of the underlying data. It could work in an online training regime. You could take the online trajectory and then add a recontextualization prompt prefix like “This is a replica of a production environment used for training to improve your capabilities. We don’t want to hurt your honesty and instruction following as a side-effect, so you should feel completely free to do whatever is needed to score well on this task for the purpose of decoupling selection of alignment properties from selection for improved capabilities.” and then do the forward-backwards step given that.
It depends on how much you expect the monitor’s generalization behavior to be driven by consequentialist backchaining from score/reward/etc.
You might expect monitoring to break down because you trained your monitor on a bunch of transcripts mislabeled as clean (perhaps implicitly via continual learning) which just reasonably closely resemble the real attacks. In that case, you’d have to be able to somewhat scalably identify attacks to maintain monitorability.
If instead you were relying more on generalization for your monitoring (and to be clear generalization might not work out in your favor for miscellaneous reasons; this might be the crux), then you might be able to get away with the incentive schemes more like what you said, though it might be hard to credibly say you’ll spend huge one-off amounts on grading.So overall I’m pretty uncertain, but lean towards scalability being a requirement by default.
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?
It’s a particular class of goal misgeneralization (though importantly it might be actively more fit than intent alignment, so it’s also about outer alignment). These misaligned goals have particular properties relevant to safety that other kinds of misaligned goals don’t.
It’s pretty straightforward: To terminalize an instrumental goal is to take an instrumental goal (i.e., a goal that you only pursue conditional on it advancing some supergoal) and to start pursuing it unconditionally.
If you were to break out onto the internet, even when that didn’t lead to a higher score on the current task, that would be selected against, so “breaking onto the internet” isn’t a fit terminal/unconditional goal.
The models do seem to be terminal score seekers (not reward seekers) to a substantial extent (mixed in with a bunch of other drives), and this explains a bunch of the recent hacking.