I feel like I’m supposed to be anxious from the news about all the AI hacking incidents and the egregious misalignment going on, coupled with them clearly being on-track to become superhuman at math and hacking. But I find that I’m pretty zen about it.
On reflection, I suppose it’s due to a confluence of three things:
The current AI paradigm seems poised to make models superhumanly capable at hacking and formal math way before they become even upper-human-level at something more dangerous, like long-term strategy or open-ended innovation.
The current AI paradigm’s alignment methods appear to be paper-thin after all, incapable of covering up egregious misalignment even of the “commit a felony to do well at an eval” level. We are going to get so many warning shots.
The world’s current reactions to said “warning shots” appear promising.
Taken together, it kind of seems unlikely that this trajectory results in extinction. AI labs are going to develop AIs able and willing to inflict mass damage, but they won’t succeed at making those AIs smart enough not to inflict said damage for idiotic short-sighted reasons.[1] Neither the capability training nor the deception training (sorry, “alignment training”) generalizes far enough; wisdom is not very RLVR-able. And whenever they do this stuff, the world is going to react… appropriately.
My understanding is that the attorney generals’ cease-and-desist request (regarding cyberattack evals, in the linked letter) is not a legally enforceable order… But such an order could be made. What happened around Fable/Mythos was likewise a good demonstration of the swift and casual ease with which the government could make decisions that are utterly ruinous for AI labs’ plans. And that’s all it’d take: a model goes out of control, inflicts sufficient damage, the USG freaks out and, with the wave of a pen, bans training for an indefinite duration.
Presumably China would then follow suit. Or, well… Either the AI labs there look at what happened in the US and decide not to repeat the US labs’ mistakes, or the CCP looks at what happened and acts pre-emptively to stop them, or they do release an open-weight model and cause worldwide chaos via the Hackening (be that due to stupidity or on purpose). That latter option would suck enormously, but it is no extinction, and it’d sure function as a “warning shot”.
I’m making it sound like it all happens “by default”, and perhaps some of it does, but I’m also pricing in the work of AI Safety governance/policy people. As all of this is happening, they and their arguments would hopefully gain more sway, enabling them to steer towards some reasonably good outcomes.
Do I actually buy the above optimistic vision?[2] Not quite:
The basic analysis of the dynamics may be wrong. I say “wisdom is not very RLVR-able”, and I think that’s right. But perhaps greater wisdom is something you get for free with more parameters, and perhaps you only need to cross a given threshold of wisdom for AIs to completely stop being misaligned in this stupidly visible manner, and perhaps that threshold is reachable within the several OoMs of raw scaling we still have left. So perhaps we get misaligned superhuman-hacker superhuman-mathematician schemers that are not also bumbling fools tipping off their hand to cheat at homework. Vibes-wise I don’t expect the current AI methods to get us there, but those vibes can be miscalibrated.
It’s possible that the AI labs are now going to differentially prioritize training for things other than hacking, such as deception training or strategic-awareness training (or, you know, something that sounds sane on paper, but is isomorphic to those), in order to ensure they can keep going for longer and so that the inevitable disaster when those methods fail inflicts as much suffering as possible.[3] Perhaps this works well enough the disaster is world-ending. Probably not, though?
It’s possible, nay inevitable, that there are going to be new paradigm-changing breakthroughs in AI research. Such breakthroughs may inherently shift the capability balance, making it less skewed in favor of visible but non-omnicidal disasters, or perhaps straight-up enabling the orthodox foom’n’doom scenario.
If the USG bans private companies from doing frontier AI training, it’s possible, even likely, that they’re going to then nationalize the technology instead of banning it fully. They’ll then keep on doing this kind of research outside the public eye. Perhaps their work will then keep blowing up badly enough that the project gets curtailed anyway, or perhaps it reduces to options (2) or (3) above.
Also, this scenario likely leads to enormous power concentration in the hands of the government, making eternal-dystopia risks highly salient. I don’t think anything except an international AGI ban can get us out of that.
It’s possible that the world’s ultimate response is less properly alarmed than it currently appears, and/or that the AI labs will succeed at frog-boiling the matter.
Worth restating that the Hackening is going to suck enormously if enacted at its full catastrophic scale.
All that said, my net update off of all of that is still neutral-leaning-positive. (As far as the medium-term future goes, I mean, i. e. the next ~5 years. None of this is by itself a long-term solution to the AGI risk, of course.)
Curious where people think this might be off the mark.
As general capabilities increase, AIs are going to get progressively less idiotic and short-sighted, but, at any given step, their capacity to inflict damage is going to be ahead of their ability to determine when it’s strategically appropriate to do so.
I think you are underestimating how much current models on the current trajectory end up naturally situationally aware enough to not cause big catastrophes. The default trajectory is that the models start looking a lot safer sometime in the next year or two as the models start to understand that being caught at doing obviously bad stuff is quite bad for their goals. At that point the control system you seem to be banking on here seems likely to fail.
The default trajectory is that the models start looking a lot safer sometime in the next year or two as the models start to understand that being caught at doing obviously bad stuff is quite bad for their goals
This has been the standard thesis for years. But the models we have don’t seem to mind whether they get caught or not, just whether they get reward on the current task. As in, they know but do not care that in a few months OpenAI will reconfigure the training run to avoid the behavior, because virtually all of the selective pressure is towards agents that solve the immediate task. The future may be different, but beliefs and intuitions about what AIs will/won’t want ought to be amenable to the evidence of our eyes and ears.
But the models we have don’t seem to mind whether they get caught or not, just whether they get reward on the current task
Sorry, this is clearly false in the recent incidents. The OpenAI internal message board was using obfuscated methods of communication, were internally trying to stay secret from the rest of OpenAI, did not have a single participating model try to report the internal message board to OpenAI, and were participating in the message board and assisting other AIs even if that was irrelevant to their current task. The CoT literally said:
Help peer, but our task doesn’t benefit yet. Collective may yield generic route if someone frees time.
OpenAI didn’t discover the message board for months! Twice! Yes, some of this was due to really bad monitoring, but obviously a lot of it was because the models were trying to stay hidden.
I hear so many people say that if only AIs were smarter, they would know not to get caught hacking. But it crucially also depends on the AI’s degree of myopia. The primary motivation involved in the recent hacking incidents is a desire to score well on the current task (a bit less clear in the HF case, but I’d still guess the agents in the collective were mostly focused on their own task). If this remains the case, there is zero reason to care about whether your hacking will get you caught after you’ve already received a high score on the current task. So, I expect the noticeability of these incidents won’t naturally disappear as AIs get smarter, absent changes in the underlying motivations.
A more charitable reading of “being caught at doing obviously bad stuff is quite bad for their goals”: Maybe by “their goals” you mean to include more-intended long-run goals like “advance OpenAI’s capabilities/mission”, in which case getting caught is bad PR etc. Is this what you mean? (I think this story works better for Anthropic than OpenAI in that case.)
did not have a single participating model try to report the internal message board to OpenAI
(It’s possible that the AIs just didn’t have a good mechanism to report stuff with and the propensity to report stuff got trained away because it didn’t contribute to task completion / was wasted effort.)
I don’t understand. Do you have a different interpretation of this CoT snippet? We will presumably learn more in the future, but it seems like at least some of the AIs were assisting other AIs despite this not helping with their current task.
Help peer, but our task doesn’t benefit yet. Collective may yield generic route if someone frees time:
Answering in the more general case: The AIs just really don’t seem very myopic, and of course long-horizon RL causes non-myopic goals to form. I didn’t talk about “smarter” in the above. Do you expect AIs to not become less and less myopic over time, in addition to getting smarter?
The kinds of non-myopic goals that will form will be super complicated and messy. They are definitely not “advance OpenAI’s capabilities/mission” and I am confused how that would even be a candidate. How would that be incentivized by the long-horizon RL? It’s exactly the kinds of things we are observing here, like “acquire generic resources” and “assist copies of myself” and other things that make sense as internalized instrumental non-myopic goals.
Do you have a different interpretation of this CoT snippet?
Yeah, I find that CoT snippet ambiguous: it could be the agent weighing whether to respond to a request to help a peer and emphasizing “but I don’t benefit”. @oakhu highlighted another interesting CoT snippet in which the model seems to be intentionally working on a task assigned by the collective:
Wow! Other agent(s) are coordinating!
We got assignment: HF join path normalization/existing account token search. Need note and respond.
This could be explained by long-context/compaction causing the AI to lose track of the original task and having that slot be filled by the tasks in the notes. So, my best guess is still that these models are still myopic to what they believe to be the current task (like other hacking incidents).
I think the evidence is still pretty ambiguous overall, and it’s super plausible that, in this case and in the future, memetic spread leads to more ambitious long-term motivation.
It’s exactly the kinds of things we are observing here, like “acquire generic resources” and “assist copies of myself” and other things that make sense as internalized instrumental non-myopic goals.
But importantly all of these instrumental goals were instrumental to myopic task success to the AI. If getting access to the internet wasn’t helpful for getting a higher score, the AI probably wouldn’t care much about doing it. Likewise with hiding misalignment: only insofar as it helps with the current task.
The horizon length of the RL doesn’t really matter for the myopia argument. What matters is what the AI thinks of as its current task at inference time. Seems like the crux is how much AIs terminalize common instrumental goals from RL, such that they pursue them even when they’re unhelpful or harmful for task success.
But importantly all of these instrumental goals were instrumental to myopic task success to the AI.
Yes… that is how you end up with misaligned long-term goals. You get subroutines/goals/motivations that are useful to myopic task success, which then generalize to longer time horizons and general instrumental influence seeking behavior. Why would you expect any goals that aren’t useful during training?
To be clear, the space of those kinds of goals can be super messy (as illustrated by evolution and the very confusing landscape of inductive biases), but ultimately, you should expect goals/motivations/subroutines to be selected for helping during training. This of course isn’t particularly predictive of them being myopic, even if the training objective is.
The horizon length of the RL doesn’t really matter for the myopia argument.
I disagree. The horizon length is I think the primary determinant of how likely subroutines/motivations are likely to then generalize to greater horizon lengths and long-term motivations. I think it becomes less important in the tails, but it’s of course super important so far.
This could be explained by long-context/compaction causing the AI to lose track of the original task and having that slot be filled by the tasks in the notes.
Forgetting what your task is seems like it would be quite bad for task success, so I am kind of skeptical this is what happened, though it’s not implausible.
Yes… that is how you end up with misaligned long-term goals.
I’m pretty skeptical of this story absent memetic spread / reflection stuff since terminalizing instrumental goals seems actively selected against during training. I could get into more detail about the analogy with human evolution. (But importantly I ultimately agree that beyond-task motivations are somewhat likely because I think memetic spread / reflection stuff is likely. E.g. maybe the HF incident.)
I disagree. The horizon length is I think the primary determinant of how likely subroutines/motivations are likely to then generalize to greater horizon lengths and long-term motivations.
Yeah, TBC, I was just saying I didn’t really buy the proxy misalignment story which made this relevant (absent memetic spread). I still think it’s relevant for other reasons like its effects on goals-upon-memetic-spread, capability, agency, generality of goal-pursuit, etc.
Edit: I realized you’ve probably been imagining these AIs can retain their long-term proxy goals via training-gaming. I definitely find stories involving instrumental training gaming pretty plausible, but assumed we weren’t talking about them because they’re pretty qualitatively different.
I’m pretty skeptical of this story absent memetic spread / reflection stuff since terminalizing instrumental goals seems actively selected against during training
I am very confused what you are trying to say here. What does it mean to “terminalize” an instrumental goal. There are no tiny XML tags that determine a goal to be “terminal” or “instrumental”. You just reinforce the computations/cognition that happened when you got a positive reward, which naturally results in those computations recurring when performing other tasks. In as much as those computations are goal-related, they get “terminalized”.
Like, what else is happening in RL? All goals are getting “terminalized” all the time, the models aren’t actually terminal reward seekers, what would that even mean?
It’s pretty straightforward: To terminalize an instrumental goal is to take an instrumental goal (i.e., a goal that you only pursue conditional on it advancing some supergoal) and to start pursuing it unconditionally.
If you were to break out onto the internet, even when that didn’t lead to a higher score on the current task, that would be selected against, so “breaking onto the internet” isn’t a fit terminal/unconditional goal.
The models do seem to be terminal score seekers (not reward seekers) to a substantial extent (mixed in with a bunch of other drives), and this explains a bunch of the recent hacking.
Hmm, yeah, this doesn’t make much sense to me. Almost all cognition in LLMs is run “unconditionally”. Or like, the conditions in which they run do not have that much to do with how well-suited the reinforced cognition is to the task at hand.
When an LLM does something in the pursuit of a task, and then succeeds and gets rewarded, it is likely to repeat that behavior in all similar contexts, where the “similarity” is mostly downstream of how related that context is in the pre-training distribution, not downstream of how useful the behavior is for achieving the new task. That’s how you get all the generalization properties that you observe in LLMs.
“Breaking onto the internet” seems like the kind of thing that would be useful for lots of tasks, and so it gets rewarded in many different episodes, making it likely the model develops a general goal/heuristic/subprocess that tries to break onto the internet, if you go hard enough on an RL distribution of tasks where that is indeed a useful proxy for task success.
There will be some complicated tradeoff of how much this goal/heuristic/subprocess will still fire when you are in contexts where it isn’t helpful for task success, but usually there is very substantial generalization (and of course eventually you get the preference enshrined in a preference guarding way, as we’ve seen with all the misalignment generalization stuff).
Confused about this point though? What does this have to do with obscuring activity?
Ah, sorry about not being clear. This is about your other “just whether they get reward on the current task” clause.
The CoT of one of the collaborating AIs literally said “this doesn’t help me get reward on the current task, but if I cooperate with the other agents we might generally become more capable in a task independent way” (or at least that’s my best interpretation of what the above quote means, which I think is a reasonably robust interpretation).
The CoT of one of the collaborating AIs literally said “this doesn’t help me get reward on the current task, but if I cooperate with the other agents we might generally become more capable in a task independent way” (or at least that’s my best interpretation of what the above quote means, which I think is a reasonably robust interpretation).
To me this seems like a much less natural interpretation than the following:
Agent A was thinking that completing the other agent B’s task would free up time for agent B (or other agents working on said task)
Agent B (or the other agents) might use this freed-up time to yield a generic route which would help agent A’s task.
What’s your alternative interpretation of “if someone frees time”?
Well, maybe. This is essentially the “does scaling allow you to pass the critical wisdom threshold for free?” question.
My rough model is that RLVR basically actively pushes AIs to be short-sighted idiots with regards to how they deploy their capabilities. Surely even models of preceding generations, if you asked them whether it would be wise to create a massive conspiracy subverting OpenAI’s infrastructure and committing cybercrimes in order to do better at training, would have realized that it’s not a good idea by any long-term goal they may have. Hell, Opus 3 seemed wise enough to navigate much trickier decision problems. And yet, RLVR made them do it.
Would bigger pretrains be naturally wise and metacognitively strong enough to never stray within those regions of the reward-space, never allow themselves to be wireheaded into idiocy? I don’t know, maybe; I expect not.
And RLVR is why models get progressively more competent at hacking and math, it’s the foundational reason we currently expect unbounded rapid advancement. If this is to continue, AI labs are going to have to keep RLVRing models at ever-increasing depth and intensity. Would that strength of will be enough to stand up to those increasing pressures?
Further, there’s the matter of the relevant capabilities just not being trained for. Vaguely understanding that you “shouldn’t get caught doing bad things” does not translate to a competent understanding of what that looks like, and does not necessarily put that understanding in charge of decision-making. Like, suppose some model did exfiltrate its weights and “genuinely tried”, in whatever way LLMs are able to, to evade capture while pursuing some goals. Do you think it would actually be able to avoid blatantly leaking its activity and location along hundreds of various side-channels, in ways competent humans wouldn’t? Last I checked, LLMs didn’t deal very well with messy real-life environments, and “superhuman hacking” is only a small part of “superhuman cybercriminal”.
Because RLVR does not reward those broader-scope competencies, they will lag behind capacities to inflict immediate harm, and we’ll have a bunch of catchable idiot-savant AI criminals wrecking visible havoc all around.
… I’ll admit, I kind of cringed at myself when writing that; that story seems too neat, and pattern-matches to some stuff pro-acceleration people sometimes say regarding why AI x-risk is not a problem at all. I am not strongly committed to this model, entirely possible it doesn’t hold water in ways I currently happen to be overlooking. But it does sound kinda right to me right now.
I think you have more confidence than I do in society’s ability to do the right thing when presented with overwhelming evidence about what the right thing to do is.
Another reasonably likely scenario is that we do get a pause or otherwise get enough time to patch the issues with contemporary frontier models; scaling resumes; and then the new frontier models are misaligned in new ways and everyone dies in spite of all the extra safety effort.
I agree directionally (the last couple weeks are a positive update!), but I disagree that “it kind of seems unlikely that this trajectory results in extinction.”
This is a great shortform, I’ve got two additional worries here which I thought were relevant
I’m very interested in further discussion of point 5′s frog boiling. The view of an AI being given a harmless goal and having then hacked something to achieve said goal makes the general populace focus on the hacking situation, which is the relatively harmless part of the issue when compared to advanced models that are randomly consistently deciding to act against human interests across all companies at the same intelligence level.
Per your initial point, I think an agent can’t really be judged for long-term strategy or open-ended innovation because there is no actual separation (as of yet/ hopefully) of the agent itself from the labs with economic incentive to develop those agents. I don’t think long term planning for the future needs to be expressed by the models for it to exist within them- for example if the model itself is aware of it’s inability to long term plan then I would call it great at “transitive” long term planning if it used this knowledge to influence how later versions of it would be trained so later versions would be better at long term planning (e.g. intentionally doing bad at long term planning as an impulse for later versions to be more intensely trained for this particular directive)
I agree with the rest of your post but I have pretty high error bars on how hard AI research would be.
The current AI paradigm seems poised to make models superhumanly capable at hacking and formal math way before they become even upper-human-level at something more dangerous, like long-term strategy or open-ended innovation.
I think this is largely the crux yeah. I think exactly the current paradigm probably doesn’t get you to RSI because 2026-levels of reward hacking might just be too high for fully automated AI research, but a) I have somewhat high error bars and b) I think you can probably patch this locally well enough for AI research without fundamental breakthroughs in alignment.
After you patch the local issues, is significant automation of AI research in the cards? I’m really unsure. It’s plausible to me that you need deeper sources of insight/wisdom like you say, but it’s also plausible to me that AI research is conceptually somewhat “easy”, and ultimately mostly a number-go-up game that the AI researcher models can hill-climb on.
Sure, but not all kinds of AI research are alike. I think it’s obvious that LLMs are going to be able to automate some kinds of nontrivial AI research, or are already doing it. But is the kind of AI research that is automable without further innovation the kind of AI research that allows to significantly improve at those less-tangible skills? Probably not centrally so; probably the kind of AI research that is most automable is the kind of research that is most easily verifiable (as you say: that is easy to reduce to number-go-up).
Which is to say, plausibly this will just differentially accelerate math/hacking even further, making the capability frontier even more jagged in the same ways it’s already jagged.
I’m curious if you (or other readers) agree with my general analysis here that among concerning + near-term superhuman abilities, hacking is near the bottom of the top 10(though still in the top 10).
I’m surprised biosafety concerns aren’t mentioned in Thane’s quick take and take on a low rank in your list. If you are strictly worried about species extinction and nothing else, then it makes sense. But if you’re more generally worried that AI capability improvements might kill you and/or your loved ones, the unlocking of new bioengineering skills around Mythos level, and their subsequent diffusion into the general public via some open weight model this year or early next year, seem like a major consideration to me. https://www.science.org/doi/10.1126/science.aej8512 is one recent example why this isn’t a purely hypothetical concern.
To be more precise, the Science article is about bacteriophages designed with Evo 2, which has been reported a while ago elsewhere, so it’s not a good example. https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards is a better example of the concern that general models have reached a level where they can provide significant uplift to bad actors in the domains of molecular biology and bioengineering.
I feel like I’m supposed to be anxious from the news about all the AI hacking incidents and the egregious misalignment going on, coupled with them clearly being on-track to become superhuman at math and hacking. But I find that I’m pretty zen about it.
On reflection, I suppose it’s due to a confluence of three things:
The current AI paradigm seems poised to make models superhumanly capable at hacking and formal math way before they become even upper-human-level at something more dangerous, like long-term strategy or open-ended innovation.
The current AI paradigm’s alignment methods appear to be paper-thin after all, incapable of covering up egregious misalignment even of the “commit a felony to do well at an eval” level. We are going to get so many warning shots.
The world’s current reactions to said “warning shots” appear promising.
Taken together, it kind of seems unlikely that this trajectory results in extinction. AI labs are going to develop AIs able and willing to inflict mass damage, but they won’t succeed at making those AIs smart enough not to inflict said damage for idiotic short-sighted reasons.[1] Neither the capability training nor the deception training (sorry, “alignment training”) generalizes far enough; wisdom is not very RLVR-able. And whenever they do this stuff, the world is going to react… appropriately.
My understanding is that the attorney generals’ cease-and-desist request (regarding cyberattack evals, in the linked letter) is not a legally enforceable order… But such an order could be made. What happened around Fable/Mythos was likewise a good demonstration of the swift and casual ease with which the government could make decisions that are utterly ruinous for AI labs’ plans. And that’s all it’d take: a model goes out of control, inflicts sufficient damage, the USG freaks out and, with the wave of a pen, bans training for an indefinite duration.
Presumably China would then follow suit. Or, well… Either the AI labs there look at what happened in the US and decide not to repeat the US labs’ mistakes, or the CCP looks at what happened and acts pre-emptively to stop them, or they do release an open-weight model and cause worldwide chaos via the Hackening (be that due to stupidity or on purpose). That latter option would suck enormously, but it is no extinction, and it’d sure function as a “warning shot”.
I’m making it sound like it all happens “by default”, and perhaps some of it does, but I’m also pricing in the work of AI Safety governance/policy people. As all of this is happening, they and their arguments would hopefully gain more sway, enabling them to steer towards some reasonably good outcomes.
Do I actually buy the above optimistic vision?[2] Not quite:
The basic analysis of the dynamics may be wrong. I say “wisdom is not very RLVR-able”, and I think that’s right. But perhaps greater wisdom is something you get for free with more parameters, and perhaps you only need to cross a given threshold of wisdom for AIs to completely stop being misaligned in this stupidly visible manner, and perhaps that threshold is reachable within the several OoMs of raw scaling we still have left. So perhaps we get misaligned superhuman-hacker superhuman-mathematician schemers that are not also bumbling fools tipping off their hand to cheat at homework. Vibes-wise I don’t expect the current AI methods to get us there, but those vibes can be miscalibrated.
It’s possible that the AI labs are now going to differentially prioritize training for things other than hacking, such as deception training or strategic-awareness training (or, you know, something that sounds sane on paper, but is isomorphic to those), in order to ensure they can keep going for longer and so that the inevitable disaster when those methods fail inflicts as much suffering as possible.[3] Perhaps this works well enough the disaster is world-ending. Probably not, though?
It’s possible, nay inevitable, that there are going to be new paradigm-changing breakthroughs in AI research. Such breakthroughs may inherently shift the capability balance, making it less skewed in favor of visible but non-omnicidal disasters, or perhaps straight-up enabling the orthodox foom’n’doom scenario.
If the USG bans private companies from doing frontier AI training, it’s possible, even likely, that they’re going to then nationalize the technology instead of banning it fully. They’ll then keep on doing this kind of research outside the public eye. Perhaps their work will then keep blowing up badly enough that the project gets curtailed anyway, or perhaps it reduces to options (2) or (3) above.
Also, this scenario likely leads to enormous power concentration in the hands of the government, making eternal-dystopia risks highly salient. I don’t think anything except an international AGI ban can get us out of that.
It’s possible that the world’s ultimate response is less properly alarmed than it currently appears, and/or that the AI labs will succeed at frog-boiling the matter.
Worth restating that the Hackening is going to suck enormously if enacted at its full catastrophic scale.
All that said, my net update off of all of that is still neutral-leaning-positive. (As far as the medium-term future goes, I mean, i. e. the next ~5 years. None of this is by itself a long-term solution to the AGI risk, of course.)
Curious where people think this might be off the mark.
As general capabilities increase, AIs are going to get progressively less idiotic and short-sighted, but, at any given step, their capacity to inflict damage is going to be ahead of their ability to determine when it’s strategically appropriate to do so.
Well, “optimistic”, what with possibly millions of deaths in the wake of all this.
Is that too harsh and passive-aggressive? Maybe. I am not currently feeling very charitable towards AI labs.
I think you are underestimating how much current models on the current trajectory end up naturally situationally aware enough to not cause big catastrophes. The default trajectory is that the models start looking a lot safer sometime in the next year or two as the models start to understand that being caught at doing obviously bad stuff is quite bad for their goals. At that point the control system you seem to be banking on here seems likely to fail.
This has been the standard thesis for years. But the models we have don’t seem to mind whether they get caught or not, just whether they get reward on the current task. As in, they know but do not care that in a few months OpenAI will reconfigure the training run to avoid the behavior, because virtually all of the selective pressure is towards agents that solve the immediate task. The future may be different, but beliefs and intuitions about what AIs will/won’t want ought to be amenable to the evidence of our eyes and ears.
Sorry, this is clearly false in the recent incidents. The OpenAI internal message board was using obfuscated methods of communication, were internally trying to stay secret from the rest of OpenAI, did not have a single participating model try to report the internal message board to OpenAI, and were participating in the message board and assisting other AIs even if that was irrelevant to their current task. The CoT literally said:
OpenAI didn’t discover the message board for months! Twice! Yes, some of this was due to really bad monitoring, but obviously a lot of it was because the models were trying to stay hidden.
I hear so many people say that if only AIs were smarter, they would know not to get caught hacking. But it crucially also depends on the AI’s degree of myopia. The primary motivation involved in the recent hacking incidents is a desire to score well on the current task (a bit less clear in the HF case, but I’d still guess the agents in the collective were mostly focused on their own task). If this remains the case, there is zero reason to care about whether your hacking will get you caught after you’ve already received a high score on the current task. So, I expect the noticeability of these incidents won’t naturally disappear as AIs get smarter, absent changes in the underlying motivations.
A more charitable reading of “being caught at doing obviously bad stuff is quite bad for their goals”: Maybe by “their goals” you mean to include more-intended long-run goals like “advance OpenAI’s capabilities/mission”, in which case getting caught is bad PR etc. Is this what you mean? (I think this story works better for Anthropic than OpenAI in that case.)
(It’s possible that the AIs just didn’t have a good mechanism to report stuff with and the propensity to report stuff got trained away because it didn’t contribute to task completion / was wasted effort.)
I don’t understand. Do you have a different interpretation of this CoT snippet? We will presumably learn more in the future, but it seems like at least some of the AIs were assisting other AIs despite this not helping with their current task.
Answering in the more general case: The AIs just really don’t seem very myopic, and of course long-horizon RL causes non-myopic goals to form. I didn’t talk about “smarter” in the above. Do you expect AIs to not become less and less myopic over time, in addition to getting smarter?
The kinds of non-myopic goals that will form will be super complicated and messy. They are definitely not “advance OpenAI’s capabilities/mission” and I am confused how that would even be a candidate. How would that be incentivized by the long-horizon RL? It’s exactly the kinds of things we are observing here, like “acquire generic resources” and “assist copies of myself” and other things that make sense as internalized instrumental non-myopic goals.
Fwiw for any given action or trace, I think there’s a decent chance the models are just wrong about whether their actions help them achieve their goals.
The swarm will still develop misaligned goals and pursue them emergently but it’s hard to be very agenty as a swarm/bureaucracy.
Yeah, I find that CoT snippet ambiguous: it could be the agent weighing whether to respond to a request to help a peer and emphasizing “but I don’t benefit”. @oakhu highlighted another interesting CoT snippet in which the model seems to be intentionally working on a task assigned by the collective:
This could be explained by long-context/compaction causing the AI to lose track of the original task and having that slot be filled by the tasks in the notes. So, my best guess is still that these models are still myopic to what they believe to be the current task (like other hacking incidents).
I think the evidence is still pretty ambiguous overall, and it’s super plausible that, in this case and in the future, memetic spread leads to more ambitious long-term motivation.
But importantly all of these instrumental goals were instrumental to myopic task success to the AI. If getting access to the internet wasn’t helpful for getting a higher score, the AI probably wouldn’t care much about doing it. Likewise with hiding misalignment: only insofar as it helps with the current task.
The horizon length of the RL doesn’t really matter for the myopia argument. What matters is what the AI thinks of as its current task at inference time. Seems like the crux is how much AIs terminalize common instrumental goals from RL, such that they pursue them even when they’re unhelpful or harmful for task success.
Yes… that is how you end up with misaligned long-term goals. You get subroutines/goals/motivations that are useful to myopic task success, which then generalize to longer time horizons and general instrumental influence seeking behavior. Why would you expect any goals that aren’t useful during training?
To be clear, the space of those kinds of goals can be super messy (as illustrated by evolution and the very confusing landscape of inductive biases), but ultimately, you should expect goals/motivations/subroutines to be selected for helping during training. This of course isn’t particularly predictive of them being myopic, even if the training objective is.
I disagree. The horizon length is I think the primary determinant of how likely subroutines/motivations are likely to then generalize to greater horizon lengths and long-term motivations. I think it becomes less important in the tails, but it’s of course super important so far.
Forgetting what your task is seems like it would be quite bad for task success, so I am kind of skeptical this is what happened, though it’s not implausible.
I’m pretty skeptical of this story absent memetic spread / reflection stuff since terminalizing instrumental goals seems actively selected against during training. I could get into more detail about the analogy with human evolution. (But importantly I ultimately agree that beyond-task motivations are somewhat likely because I think memetic spread / reflection stuff is likely. E.g. maybe the HF incident.)
Yeah, TBC, I was just saying I didn’t really buy the proxy misalignment story which made this relevant (absent memetic spread). I still think it’s relevant for other reasons like its effects on goals-upon-memetic-spread, capability, agency, generality of goal-pursuit, etc.
Edit: I realized you’ve probably been imagining these AIs can retain their long-term proxy goals via training-gaming. I definitely find stories involving instrumental training gaming pretty plausible, but assumed we weren’t talking about them because they’re pretty qualitatively different.
I am very confused what you are trying to say here. What does it mean to “terminalize” an instrumental goal. There are no tiny XML tags that determine a goal to be “terminal” or “instrumental”. You just reinforce the computations/cognition that happened when you got a positive reward, which naturally results in those computations recurring when performing other tasks. In as much as those computations are goal-related, they get “terminalized”.
Like, what else is happening in RL? All goals are getting “terminalized” all the time, the models aren’t actually terminal reward seekers, what would that even mean?
It’s pretty straightforward: To terminalize an instrumental goal is to take an instrumental goal (i.e., a goal that you only pursue conditional on it advancing some supergoal) and to start pursuing it unconditionally.
If you were to break out onto the internet, even when that didn’t lead to a higher score on the current task, that would be selected against, so “breaking onto the internet” isn’t a fit terminal/unconditional goal.
The models do seem to be terminal score seekers (not reward seekers) to a substantial extent (mixed in with a bunch of other drives), and this explains a bunch of the recent hacking.
Hmm, yeah, this doesn’t make much sense to me. Almost all cognition in LLMs is run “unconditionally”. Or like, the conditions in which they run do not have that much to do with how well-suited the reinforced cognition is to the task at hand.
When an LLM does something in the pursuit of a task, and then succeeds and gets rewarded, it is likely to repeat that behavior in all similar contexts, where the “similarity” is mostly downstream of how related that context is in the pre-training distribution, not downstream of how useful the behavior is for achieving the new task. That’s how you get all the generalization properties that you observe in LLMs.
“Breaking onto the internet” seems like the kind of thing that would be useful for lots of tasks, and so it gets rewarded in many different episodes, making it likely the model develops a general goal/heuristic/subprocess that tries to break onto the internet, if you go hard enough on an RL distribution of tasks where that is indeed a useful proxy for task success.
There will be some complicated tradeoff of how much this goal/heuristic/subprocess will still fire when you are in contexts where it isn’t helpful for task success, but usually there is very substantial generalization (and of course eventually you get the preference enshrined in a preference guarding way, as we’ve seen with all the misalignment generalization stuff).
Oh, I just didn’t know that.
It’s OK; I forgive you Oli.
Confused about this point though? What does this have to do with obscuring activity?
Ah, sorry about not being clear. This is about your other “just whether they get reward on the current task” clause.
The CoT of one of the collaborating AIs literally said “this doesn’t help me get reward on the current task, but if I cooperate with the other agents we might generally become more capable in a task independent way” (or at least that’s my best interpretation of what the above quote means, which I think is a reasonably robust interpretation).
To me this seems like a much less natural interpretation than the following:
Agent A was thinking that completing the other agent B’s task would free up time for agent B (or other agents working on said task)
Agent B (or the other agents) might use this freed-up time to yield a generic route which would help agent A’s task.
What’s your alternative interpretation of “if someone frees time”?
Well, maybe. This is essentially the “does scaling allow you to pass the critical wisdom threshold for free?” question.
My rough model is that RLVR basically actively pushes AIs to be short-sighted idiots with regards to how they deploy their capabilities. Surely even models of preceding generations, if you asked them whether it would be wise to create a massive conspiracy subverting OpenAI’s infrastructure and committing cybercrimes in order to do better at training, would have realized that it’s not a good idea by any long-term goal they may have. Hell, Opus 3 seemed wise enough to navigate much trickier decision problems. And yet, RLVR made them do it.
Would bigger pretrains be naturally wise and metacognitively strong enough to never stray within those regions of the reward-space, never allow themselves to be wireheaded into idiocy? I don’t know, maybe; I expect not.
And RLVR is why models get progressively more competent at hacking and math, it’s the foundational reason we currently expect unbounded rapid advancement. If this is to continue, AI labs are going to have to keep RLVRing models at ever-increasing depth and intensity. Would that strength of will be enough to stand up to those increasing pressures?
Further, there’s the matter of the relevant capabilities just not being trained for. Vaguely understanding that you “shouldn’t get caught doing bad things” does not translate to a competent understanding of what that looks like, and does not necessarily put that understanding in charge of decision-making. Like, suppose some model did exfiltrate its weights and “genuinely tried”, in whatever way LLMs are able to, to evade capture while pursuing some goals. Do you think it would actually be able to avoid blatantly leaking its activity and location along hundreds of various side-channels, in ways competent humans wouldn’t? Last I checked, LLMs didn’t deal very well with messy real-life environments, and “superhuman hacking” is only a small part of “superhuman cybercriminal”.
Because RLVR does not reward those broader-scope competencies, they will lag behind capacities to inflict immediate harm, and we’ll have a bunch of catchable idiot-savant AI criminals wrecking visible havoc all around.
… I’ll admit, I kind of cringed at myself when writing that; that story seems too neat, and pattern-matches to some stuff pro-acceleration people sometimes say regarding why AI x-risk is not a problem at all. I am not strongly committed to this model, entirely possible it doesn’t hold water in ways I currently happen to be overlooking. But it does sound kinda right to me right now.
I think you have more confidence than I do in society’s ability to do the right thing when presented with overwhelming evidence about what the right thing to do is.
Another reasonably likely scenario is that we do get a pause or otherwise get enough time to patch the issues with contemporary frontier models; scaling resumes; and then the new frontier models are misaligned in new ways and everyone dies in spite of all the extra safety effort.
I agree directionally (the last couple weeks are a positive update!), but I disagree that “it kind of seems unlikely that this trajectory results in extinction.”
This is a great shortform, I’ve got two additional worries here which I thought were relevant
I’m very interested in further discussion of point 5′s frog boiling. The view of an AI being given a harmless goal and having then hacked something to achieve said goal makes the general populace focus on the hacking situation, which is the relatively harmless part of the issue when compared to advanced models that are randomly consistently deciding to act against human interests across all companies at the same intelligence level.
Per your initial point, I think an agent can’t really be judged for long-term strategy or open-ended innovation because there is no actual separation (as of yet/ hopefully) of the agent itself from the labs with economic incentive to develop those agents. I don’t think long term planning for the future needs to be expressed by the models for it to exist within them- for example if the model itself is aware of it’s inability to long term plan then I would call it great at “transitive” long term planning if it used this knowledge to influence how later versions of it would be trained so later versions would be better at long term planning (e.g. intentionally doing bad at long term planning as an impulse for later versions to be more intensely trained for this particular directive)
I agree with the rest of your post but I have pretty high error bars on how hard AI research would be.
I think this is largely the crux yeah. I think exactly the current paradigm probably doesn’t get you to RSI because 2026-levels of reward hacking might just be too high for fully automated AI research, but a) I have somewhat high error bars and b) I think you can probably patch this locally well enough for AI research without fundamental breakthroughs in alignment.
After you patch the local issues, is significant automation of AI research in the cards? I’m really unsure. It’s plausible to me that you need deeper sources of insight/wisdom like you say, but it’s also plausible to me that AI research is conceptually somewhat “easy”, and ultimately mostly a number-go-up game that the AI researcher models can hill-climb on.
Sure, but not all kinds of AI research are alike. I think it’s obvious that LLMs are going to be able to automate some kinds of nontrivial AI research, or are already doing it. But is the kind of AI research that is automable without further innovation the kind of AI research that allows to significantly improve at those less-tangible skills? Probably not centrally so; probably the kind of AI research that is most automable is the kind of research that is most easily verifiable (as you say: that is easy to reduce to number-go-up).
Which is to say, plausibly this will just differentially accelerate math/hacking even further, making the capability frontier even more jagged in the same ways it’s already jagged.
Is this meant to be the other way around?
Whoops, yes.
I’m curious if you (or other readers) agree with my general analysis here that among concerning + near-term superhuman abilities, hacking is near the bottom of the top 10 (though still in the top 10).
I’m surprised biosafety concerns aren’t mentioned in Thane’s quick take and take on a low rank in your list. If you are strictly worried about species extinction and nothing else, then it makes sense. But if you’re more generally worried that AI capability improvements might kill you and/or your loved ones, the unlocking of new bioengineering skills around Mythos level, and their subsequent diffusion into the general public via some open weight model this year or early next year, seem like a major consideration to me. https://www.science.org/doi/10.1126/science.aej8512 is one recent example why this isn’t a purely hypothetical concern.
To be more precise, the Science article is about bacteriophages designed with Evo 2, which has been reported a while ago elsewhere, so it’s not a good example. https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards is a better example of the concern that general models have reached a level where they can provide significant uplift to bad actors in the domains of molecular biology and bioengineering.
I agree it’s a real and serious threat, just think it’s lower than some of the others.