I’m an AGI safety / AI alignment researcher in Boston with a particular focus on brain algorithms. Research Fellow at Astera. I’m also at: Substack, X/Twitter, Bluesky, RSS, email, and more at this link. See https://sjbyrnes.com/agi.html for a summary of my research and sorted list of writing. Physicist by training. Leave me anonymous feedback here.
Steven Byrnes
Specifically, it seems plausible/likely to me that if you take some transformerish like architecture and throw enough task based RL compute at it, the general program search finds genuinely dangerous agentic programs which could execute a takeover, whether or not they bootstrap themselves to novel architectures.
I mean, I certainly have concerns about LLM alignment (see Bonus section). But the opposite of “AI will definitely be egregiously misaligned (in the absence of new breakthrough alignment ideas)” (per Yudkowsky, Soares, me) is not “AI will definitely not be egregiously misaligned…” but rather “AI might or might not be egregiously misaligned…”. And arguing against that is much harder, because you need to argue not only that there are potential problems but that these problems are inevitable, with no possible adequate mitigations within the domain of known techniques.
So then we need to start going through possible mitigations to LLM misalignment, and whether those mitigations are durable solutions versus merely delaying the inevitable. And in near-future-LLM world, there are definitely mitigations! Like, LLMs today are obviously not being trained using best known practices for alignment. E.g. the companies keep accidentally using buggy RL environments that reward the LLM-in-training for cheating, and they keep accidentally exposing chains-of-thought to the reward functions, etc. They could fix that. And on top of that, there’s a whole world of other “obvious” mitigations that we would have to talk about. E.g. what if we do another round of DPO after the RLVR, but the DPO is a team of super-conscientious world experts spending hours judging each output? What if the humans are scrutinizing the chain-of-thought too? Or what if we simply limit the amount of RLVR, rather than scaling up RLVR training forever, and instead treat the RLVR phase as a bootstrapping step for massively-scaled-up DPO and RLAIF? How can we make the RLAIF judge models better? Etc. etc. There’s a whole argument tree here, and as of now I’m at least vaguely sympathetic to the LLM people who wind up feeling like the answer is “If we keep using the kinds of LLM training approaches that we’re using today, those future more powerful LLMs might or might not be egregiously misaligned”. (“Vaguely sympathetic” is weaker than “agree with”; I just don’t have a very strong opinion either way.)
I mean, sure you could lump them together, but I think the strategies for getting human approval do not exactly match the strategies for getting LLM approval, even if there’s some overlap.
By the way, in another comment I also suggested that we might also draw a distinction between alignment-targeted approval versus capabilities-targeted approval, regardless of whether the approval is from an human or an AI. (Again, there’s overlap, and it’s a blurry line separating them.)
The “flavors of misalignment” aren’t mutually exclusive. But for any particular kind of misalignment, you can still ask the question “Where did that come from?”, in the sense that random policies are not misaligned (merely useless), so any a-priori-unlikely recognizable behavior has to come from somewhere (cf. “the follow-the-improbability game”), and it’s probably one of the stages of training.
And I’m suggesting that for a certain recognizable kind of misaligned behavior (the kind involving pride, jealousy, trolling, etc.), the answer to “where did it come from” is the pretraining / SFT stage, as opposed to DPO or RLVR etc. Whereas other kinds of misaligned behaviors originate in different training stages.
So anyway, it’s not unavoidable in principle (e.g. you could filter all not-nice human dialogue out of the pretraining data), but yeah it is universal in LLMs to date.
Yeah diffusion LLMs are still trained by imitative learning (“the magical transmutation of observations into behavior”). So if the training data has lots of dialog text in which a crazy upset woman tries to convince a guy to leave his wife, then a diffusion LLM (just like a normal LLM) will be capable of outputting dialog text in which a crazy upset woman tries to convince a guy to leave his wife, like Bing-Sydney did.
I think that, to the extent that SFT is determining model outputs, you get an LLM that imitates whatever the SFT data is. (At least to a first approximation.)
So e.g. if you have ruthless-optimizer literal-genie RLVR traces, and you distill (do SFT on) those traces, then you can get a ruthless-optimizer literal-genie LLM, even though mathematically speaking you were doing SFT not RLVR.
do you share my intuition the pretraining/SFT category looks like the least scary one by far
Yeah seems reasonable. But the LLMs trained that way are less capable than the LLMs trained using RLVR etc. So here we are.
So far, I hope I have understood you correctly.
Yup!
How does one grow a good human?
I think human innate drives (especially social instincts) are critical. E.g. there are plenty of sociopaths who grow up in loving families.
young humans being malleable and not perfectly ruthless responders to optimization pressure
I get a lot of pushback on this, but I strongly believe that people are extraordinarily responsive to optimization pressure, pretty much every waking second of every day, no exceptions (see here). But the optimization pressure in question is internal, coming from our innate drives, which are brain signals that trigger for lots of very-not-obvious reasons, in lots of superficially-quite-different circumstances. I think it’s a very common error for people to assume that human optimization pressure is more external than it is, e.g. this handy chart I made in 2023 ↓
Might we go back to imitative learning on already-grown humans
There’s a lot to be said for imitative learning in terms of safety. But it doesn’t get you to superhuman capabilities. Indeed, we had mostly imitative learning, back in the good old pre-o1 days of mid-2024, and then the AI companies noticed that the models weren’t as capable as they wanted, so they “solved” that “problem” by doing more and more non-imitative-learning post-training.
Re-reading this, it’s possible that I should have split up “RLAIF-for-alignment” versus “RLAIF-for-capabilities” into two separate rows? (The first includes constitutional AI & deliberative alignment, whereas the second would be things like having a judge model pick holes in math proofs.) I guess there isn’t a sharp line between them, but maybe they’re different enough that we shouldn’t intuitively lump them together? Not sure.
There’s a similar loose split in RLHF/DPO (it can be geared more towards alignment versus more towards pushing the limits of capabilities, depending on what human is selecting answers and how). I did mention that one in the post, but only as an aside.
Four LLM loss functions → four flavors of LLM misalignment
I think you’re thinking about “automation” in a very different way from me. Copying from something I wrote here:
…what exactly is our “AGI Jeff Bezos” supposed to be doing at any given time?
Nobody knows!
Indeed, the fact that nobody knows is the whole point! That’s the very reason that an AGI Jeff Bezos can create so much value!
When Human Jeff Bezos started Amazon in 1994, he was obviously not handed a detailed spec for what to do in any possible situation, where following that spec would lead to the creation of a wildly successful e-commerce / cloud computing / streaming / advertising / logistics / smart speaker / Hollywood studio / etc. business. For example, in 1994, nobody, not Jeff Bezos himself, nor anyone else on Earth, knew how to run a modern cloud computing business, because indeed the very idea of “modern cloud computing business” didn’t exist yet! That business model only came to exist when Jeff Bezos (and his employees) invented it, years later.
By the same token, on any given random future day…
Our AGI Jeff Bezos will be trying to perform a task that we can’t currently imagine, using ideas and methods that don’t currently exist.
It will have an intuitive sense of what constitutes success (on this micro-task) that it learned from extensive idiosyncratic local experience, intuitions that a human would need years to replicate.
The micro-task will advance some long-term plan that neither we nor even the AGI can yet dream of.
This will be happening in the context of a broader world that may be radically different from what it is now.
And our AGI Jeff Bezos (along with other AGIs around the world) will be making these kinds of decisions at a scale and speed that makes it laughably unrealistic for humans to be keeping tabs on whether these decisions are good or bad.
Thus, if you automate all the tasks that Jeff Bezos has ever done, then you still have not automated Jeff Bezos. Indeed, I would say that you haven’t yet even begun to automate Jeff Bezos. Because the key value of Jeff Bezos is his ability to figure out how to do new things that nobody has ever done before.
For further discussion, see also: What do I mean by “Artificial General Intelligence”? E.g.: “…Many copies of one human brain design, barely changed since the African savannah, built the global economy from scratch … If you want a human to do [a new task], you don’t need to do R&D to breed a new subspecies of human!”
UPDATE 8/10: I partly retract this comment; I was talking about DPO below when I should have been talking about RLAIF instead. See Four LLM loss functions → four flavors of LLM misalignment
It’s tricky to think about this topic because we don’t know exactly how the AI companies train their LLMs. (Or at least I sure don’t.) But here’s a speculative possibility to explain your puzzle that “the RL capabilities generalize further than the RL misalignment”:
Famously (Ord, Dwarkesh, Beren, etc.), RLVR needs bootstrapping to work, because RLVR doesn’t provide many bits, and there’s no training signal if it never succeeds. The bootstrapping comes from pretraining / SFT.
Likewise, DPO-on-a-hard-task (like refactoring an entire big codebase) needs bootstrapping to work, because DPO doesn’t provide many bits, and the human training signal isn’t very helpful if it never gets anywhere remotely close to refactoring the codebase. And maybe this bootstrapping partly comes from RLVR.
I.e., maybe after RLVR, there’s another expert-DPO stage, and that latter stage gets the model from RLVR-craziness closer to all-things-considered human preferences, at least in DPO-like contexts (which would be more like deployment and less like “graded episodes”). I think this would explain why “the RL capabilities generalize further than the RL misalignment”: the final expert-DPO-on-very-hard-tasks stage can simply directly select for both the capabilities and the (at least superficial) alignment simultaneously.
Now, the mantra that “RL is evil” (Bengio) or “RL is terrifying” (me) or “I’m still scared of RL” (Paul Christiano) is a great default starting point, but it does require some caveats and footnotes. In expert-DPO, there’s a human examining and grading the gestalt of what the LLM has spit out, based on the spirit of their own human intentions. If the LLM optimizes for that, then yes you can get sycophancy or trickery, but if the graders are good and careful, it can be mitigated. The graders can also check the tool call logs for hijinks. Thus DPO is not (necessarily) “evil” in the way that RLVR is. In particular, DPO has some of the safety benefits of “MONA”, in a way that RLVR does not.
If all that is right, is it a sustainable solution to LLM alignment? I feel like my mind goes to my traditional answer: yes, if and only if LLMs never get to ASI. To the extent that this system creates good results, I think it’s load-bearing that both the RLVR and the expert-DPO final stage are mere trickles of bits, not the kind of processes that can build up giant edifices of new superhuman knowledge and capabilities at scale. After all, in the limit, DPO does break down; it relies on the human graders being capable of oversight.
(It’s also possible that the companies go back and forth between expert-DPO and RLVR during training, which would make it even more obvious that we should expect LLMs to wind up with strong behavioral divergence, with ruthlessness in graded-episode contexts and with what-the-DPO-expert-graders-would-want in other contexts.)
To add on to that:
Okay, but CoastRunners is, after all, the canonical example of weird over-optimization, the one everyone brings up. For the canonical example to be just-plain-wrong, and obviously so… well, it feels emblematic of the larger problem, to me.
I’m feeling defensive here because I have occasionally cited that example for that purpose.
There is a set of situations where the result of optimization will include Behavior X, but where the person who set up the optimization had not thought of X ahead of time and is surprised when it happens.
In some of those cases, we can all have a good laugh at how dumb the person who set up the optimization must have been, to have not realized that X would result from the optimization, because c’mon, duh, it should have been blindingly obvious from the start.
In the opposite extreme, we can all feel very sympathetic to the person not thinking of X in advance, because man, I wouldn’t have thought of X either.
(Also, lots more X’s seem obvious in hindsight than in foresight.)
Anyway, the important point is the phenomenon in which the person is surprised to discover Behavior X. And I think CoastRunners is a fine example of that happening (assuming that Dario & Jack were in fact surprised, which seems quite plausible to me, even if you want to laugh at them for that).
Also, if you’re thinking “surely smart people like Dario & Jack would have recognized in advance the obvious-to-me consequences of optimizing (blah)”, then please witness the brazen failures of very intelligent people like LeCun, Silver & Sutton, Schmidhuber, Musk, etc. to recognize in advance the obvious-to-many-people consequences of optimizing (blah), even with high stakes and even after other people have repeatedly pointed out their error. Apparently it’s just really hard for many people to think clearly in advance about the results of an optimization search.
Consider “reward hacking.” Or, ~equivalently, “specification gaming.” …Goodhart’s law … it has the same odd framing issue as the other terms: ascribing the undesired outcome to the optimizer, rather than to the optimized function, its argmax, and/or the person who came up with it to begin with. Hating the player, instead of the game.
Yeah, I like to talk about emotive conjugations: Just like “I am firm, you are obstinate”, or “I am unique, you are a weirdo”, we also have: “I find out-of-the box solutions, you engage in specification-gaming”.
…there is no reason to expect optimal score-maxxing play to look anything like play that optimizes for sequence-of-stages completion … This is true in CoastRunners, and in many arcade-style games, and also in, like… Mario, right? It’s not exactly some obscure phenomenon known only to hardcore gamers.
At the time, the talk of the town was the DeepMind Atari-playing DQN (2013, 2015), which almost used pure Atari game score as reward, except that (1) they clipped it to (+1 for any increase, −1 for any decrease, or 0), and (2) they reset the training after one death even if the game allowed multiple lives. According to some LLM I asked just now, the 2015 version fully beat Pong and Boxing, and progressed on a number of other games without getting to the end, but it also found some weird repetitive point-farming strategy in Kangaroo.
I was making a narrow point about what I mean by “LLMs are (still) mostly powered by imitative learning, not RL”.
You seem to be saying: “Nobody ever would have expected that RLVR to build a whole giant new edifice of deep interconnected knowledge into an LLM, the way AlphaZero’s MCTS+RL training did. Duh, that’s obvious and overdetermined. So why are we even talking about that?”
But to the extent that that’s true, it would make my point stronger, not weaker. Because imitative learning can (and does) build a whole giant new edifice of deep interconnected knowledge into an LLM.
It seems like you’re trying to argue about takeoff speeds instead? If so, that seems off-topic here, I think. My take on that is at Foom & Doom 1: “Brain in a box in a basement”, including the part starting at “To be clear, the resulting ASI after those 0–2 years would not be an AI that already knows everything about everything…”
Yeah I agree that people seem to treat anthropic arguments as a special category instead of one piece of evidence among many, I was arguing about that once here.
I don’t know much about GFlowNets. Anyway, that paper seems very specific to LLMs, which are off-topic for this post (see Q1 & Q7), sorry.
One line of thought I have here is, there are lots of things such a human or AGI could disagree or talk about or have an interest in, how does it pick which one? I think for the human it probably comes down to some kind of subconscious status calculation, but in either case, how does the AGI do it if it doesn’t have its own status motivations or other long-term goals?
I think we do want the AGI to have long-term goals, just not to have exclusively long-term goals, at least in the §6.2.1 approach. See my old 2021 post Consequentialism & corrigibility. Again, if you or me is the model to be inspired by, then I assume we both care about the future of life being great (long-term goal), but we both also enjoy figuring things out in the here and now, and we both also have principles that we take pride in. And for my part, I wouldn’t want to be benevolent dictator of the universe even if I could, that’s way too much responsibility, sounds terrifying.
To be clear, I’m generally expecting the process to be kinda messy, where it’s kinda hard to reason about where the AGI winds up. From my perspective, Step 1 is to have any plan at all that could plausibly work, and then we can move on to making it easier to test and de-risk the plan in advance, to the extent possible. As mentioned at the bottom, I’m still hard at work trying to get more clarity, to the extent possible, in order to make the process of AGI motivation development more predictable and legible and less messy.
I don’t think that theory [social status] explains quite as much as those people think it does
Do you have a link/explanation for this? I think this may be fairly cruxy, because I’m guessing your intuitions for truth-seeking disagreeable nerd AGI are substantially based on truth-seeking disagreeable nerd humans, so it matters what those humans’ real motivations are.
Not all in one place, but here’s some pointers (and feel free to ask follow-ups).
Let’s start with the Hansonian notion of strategic self-deception (as distinct from plain old motivated reasoning which is real and important), which I assume he got from Robert Trivers. I broadly reject that notion. E.g. here’s a footnote in my post about laughter:
Since the works of Robin Hanson are popular on this forum, I will say a bit more about where I differ from Elephant In The Brain. My biggest complaint is the part where they say:
As we mentioned earlier, people are profoundly ignorant about laughter’s meaning and purpose (at least in our default state, before learning the science). But where does this ignorance come from? Why does introspection fail us so spectacularly here?
It’s not simply because laughter is involuntary, outside our conscious control. Flinching, for example, is also involuntary, and yet we understand perfectly well why we do it: to protect ourselves from getting hit. Thus our ignorance about laughter needs further explanation.
I disagree that it “needs further explanation”. I think we start out ignorant of literally everything, until we learn it / figure it out. And I think that figuring out the evolutionary purpose of laughter is just inherently much harder than figuring out the evolutionary purpose of flinching. It’s less obvious / salient, for various reasons that I claim are pretty obvious if you think about it. I don’t think there’s any more to it than that.
I also don’t think there can be more to it than that. To explain what I mean by that, imagine if I said: “Here’s the source code for training an image-classifier ConvNet from random initialization using uncontrolled external training data. Can you please edit this source code so that the trained model winds up confused about the shape of Toyota Camry tires specifically?” The answer is: “Nope. Sorry. There is no possible edit I can make to this PyTorch source code such that that will happen.” By the same token, even if, as that book argues, there is a strong evolutionary pressure to make humans specifically confused about the evolutionary purpose of laughter, I don’t think there is any possible genetic change that would make that happen. Related discussion here.
Up a level, this strategic self-deception idea comes out of the “evolved modularity” framework in evolutionary psychology, and I reject that whole broader framework as well. See §1.1 of “My take on Jacob Cannell’s take on AGI safety” for the background, and “Learning from scratch” in the brain for why I don’t buy it (kinda related to the thing above about PyTorch).
For related reasons, I reject the idea that “status” (per se) could possibly be an innate goal. It’s just too abstract. Copying from here:
Explaining how human social instincts work is tricky mainly because of the “symbol grounding problem”. In brief, everything we know—all the interlinked concepts that constitute our understanding of the world and ourselves—is created “from scratch” in the cortex by a learning algorithm, and thus winds up in the form of a zillion unlabeled data entries like “pattern 387294 implies pattern 579823 with confidence 0.184”, or whatever. Yet certain activation states of these unlabeled entries—e.g., the activation state that encodes the fact that Jun just told me that Xiu thinks I’m cute—need to somehow trigger social instincts in the Steering Subsystem. So there must be some way that the brain can “ground” these unlabeled learned concepts.
Now, it’s not that there’s no way to solve the symbol grounding problem to make someone want social status—indeed, that obviously happens, and in my post Neuroscience of human social instincts: a sketch, I attempt to explain how. It’s that there’s no way to solve the symbol grounding problem to make someone want social status specifically. Realistically, the genetic mechanism is just not gonna be that specific. (More on which shortly.)
So here’s where we’re at so far: (1) without Trivers-style self-deception, we have a harder time explaining away introspective reports like “That’s not status-seeking because I’m not trying to impress anyone” (we can still try to explain it away, but it’s harder); and (2) the idea that “status” could have a special place as an innate end-goal is implausible anyway.
That brings us to my Social drives 2: “Approval Reward”, from norm-enforcement to status-seeking, which is a follow-up to Neuroscience of human social instincts: a sketch where I attempt to connect the neuroscience to everyday life:
That post’s §2 is the part that’s closest to status-seeking: people are motivated to have actual interactions where an actual person has positive associations with you. But even in this case, status-seeking is just one of several consequences of the same innate drive. The others are credit-seeking / blame-avoidance, and norm-following / norm-enforcement. I understand that you can try to unify these by hypothesizing that (say) credit-seeking is a means-to-an-end for achieving social status, but that’s just not how it works in my neuroscience model: credit-seeking / blame avoidance, status-seeking, and norm-enforcement all emerge in the same way, at the same level, via the same mechanism.
And then that post’s §3 gets even more distant from the conventional notion of status-seeking, by analyzing how we can feel pride in ourselves and our actions, and how this is another direct consequence of the same innate drive, and not a secret means-to-an-end to achieving social status.
And indeed, it’s easy to come up with cases where pride vs future-status come apart, and where people follow the former over the latter. E.g. people sometimes stand up for principles even if they think everyone will scorn them for it, because they’re following their own moral compass. Certainly the moral compass has something to do with what other people have said and thought over the course of the person’s life, but the relation can be quite indirect (e.g. people may care about how a cartoon character would judge their behavior), and not well-described as “trying to wind up with high social status”, consciously or unconsciously.
CoT that’s legible isn’t good evidence that LLMs are not wielding (or could not wield) novel latent abstractions/reasoning primitives/control heuristics.
I feel like you’re saying: LLMs can wield more than zero novel latent abstractions/reasoning primitives/control heuristics while CoT remains generally legible.
Whereas what I’m saying is: LLMs cannot be totally transformed by RLVR while CoT remains generally legible.
These aren’t contradictory. I think both are true.
As an example, think about humans learning things, like a teen going from her first number theory class as a teen at time 0, to deeply understanding very advanced math (e.g. the Langlands program) as an adult at time T = many years later. It’s an arduous and time-consuming process. And her notes at time T would be deeply, deeply inscrutable from the perspective of her former teen self at time 0—even the notes that are in the form of words rather than symbols.
This suggests that the delta between pre-RLVR vs post-RLVR LLMs is much less of a wrenching change than the delta between the teen at time 0 vs the now-adult mathematician at time T.
Novelty builds heavily on existing frameworks and often involves a subtle reframe, reconfiguration, or saliencing of a known/slightly modified concept in a new setting. Insight is about discovering relevance, but what’s available to be relevant is often familiar primitives.
Mathematical work is a great example here, since mathematical discoveries very often have a subtle kernel of insight, a small new idea, that reconfigures existing concepts around a problem in a significant way to illuminate something unknown. The reconfiguration is almost all in terms of familiar stuff—the moving pieces don’t change much. In the case of LLMs, that’s the content imitative learning provides.
I think instead of “novelty” here you should have said “a sufficiently small increment of novelty”.
If we instead consider mathematics as a collective human enterprise, it went from “number theory doesn’t exist at all” to the Langlands program, over the course of 200 years.
Maybe it sounds absurd for me to compare what one RLVR training can do, versus the whole edifice of ideas painstakingly built by the mathematics community over the course of 200 years. But it’s not absurd: AlphaZero really did blow past the whole edifice of ideas painstakingly built over the course of centuries by the chess and go communities in its 72-hour training runs. So the idea of building real new knowledge at a massive scale through RL is not absurd on its face. I’m just saying: RLVR-on-LLMs is not doing anything like that, at least not today. Rather, the LLM approach is to use imitative learning to suck in the whole edifice of ideas painstakingly built by humans, and then tweak it a bit at the end via RLVR.
Relatedly, when mathematicians study the recent LLM-generated math results, they’re not “learning something new” in a way that’s analogous to that teen spending years poring over her math textbooks. Rather they’re “learning something new” in a way that’s analogous to some guy telling me what his name is. If the mathematicians already have all the right background knowledge, they can quickly understand the solution within their existing conceptual frameworks. Or if they don’t already have all the right background knowledge, they can read human-created textbooks to get it. (I’m mainly thinking of the unit distance conjecture; in other examples that I looked into like the Jacobian conjecture, IIUC, the CoTs weren’t released, so mathematicians are still be a bit puzzled about how the LLMs came upon the answer.)
…LLMs too, seem to perform much of their computation through nonlinguistic representations and compress the actionable results into language, mostly because of the structural thing that language is the medium through which they maintain serial state and communicate.
This is an argument against
if you create AI capabilities via imitative learning, you get models that follow the human distribution of outputs
I don’t think it’s an argument against that, or sorry if I’m misunderstanding. You’re saying that LLMs may be computing their outputs in a different way from humans. Fine. But their outputs are still following the human distribution of outputs. “Following the human distribution of outputs” is just another way to say “low perplexity”, right?
(I will concede that the phrase “the human distribution of outputs” is a bit misleading on various other grounds, e.g. LLMs may generalize OOD in a different way from humans; not all imitative learning training tokens are created by humans; etc.)
Oh, I wasn’t talking about changing the LLM learning algorithm, training approach, etc. Of course if you change any of those things, then you can get different results for both capabilities and alignment. I meant, like, I don’t expect LLMs as trained today to pass some threshold beyond which they will logically reason themselves into being egregiously misaligned.