I’m an AGI safety / AI alignment researcher in Boston with a particular focus on brain algorithms. Research Fellow at Astera. I’m also at: Substack, X/Twitter, Bluesky, RSS, email, and more at this link. See https://sjbyrnes.com/agi.html for a summary of my research and sorted list of writing. Physicist by training. Leave me anonymous feedback here.
Steven Byrnes
Do you just mean that none of those other drives are necessary to make the AGI safe, or do you mean that they should all be omitted from the (first) AGI?
I meant the latter, although I don’t feel very strongly.
In fact, hmm, if we’re making an AGI singleton sovereign (either on purpose [cf. §6.2.3] or by accident), then anything we leave out of its value system will quite possibly get left out of the whole future forever (unless it defers to humans), so if the AGI feels no connection to laughter and love and so on, that’s at least a potential problem. Not sure how it balances against other things.
But yeah, if it’s just one short-lived AGI, I don’t think its subjective experience is super-high on my list of concerns, as an end in itself. I’m much more concerned about its decisions that impact the whole future. (Of course, its subjective experience is relevant to how it makes decisions.)
how do you think the AGI should feel about images of/representing humans? e.g. photos, drawings, and cartoons? Should the AGI’s innate drives trigger on them or not? Should it treat photorealistic images differently from cartoons?
I don’t want an AGI to intrinsically care about the welfare of a photo of a human. Do you? We can talk about why it might be tricky to make that happen, but the ideal goal seems clear.
(I recall that human empathy can trigger on very simple drawings, but I don’t know how much of that is directly from human innate drives and how much is generalized.)
Are you thinking of the Heider & Simmel 1944 animations by any chance?
even if the bad AGI has the option to wipe out humanity with a surprise attack with bioweapons etc, if the good AGI would be able to survive, counterattack, and destroy the bad AGI afterwards, then the bad AGI might not want to attack in the first place
A second-strike-capable “good” AGI might be an adequate solution to the problem of “bad” rogue AGI, but I think that would wind up involving the “good” AGI doing pretty scary and aggressive things: probably the “good” AGI needs to get itself running on a majority of compute on Earth, harden that compute against cyber and physical attack, make sure both those chips and those defense systems are not reliant on human infrastructure (e.g. robots defending the chips or whatever), do as much recursive self-improvement as possible, and so on. Otherwise a power-seeking misaligned AGI would be able to crush the “good” AGI along with the humans.
So I see that as a plausible solution to the “strategic” problem of §6.1, but like all the other solutions I know of, it would (1) require AGI superpowers to make it happen, and (2) have aspects that I think would strike people as quite scary and unpalatable. You’re welcome to disagree with any of that, of course.
Thanks!
I listened to a Spencer Greenberg podcast on OCPD a few weeks ago and it seems pretty consistent with (and complementary to) what you wrote.
Of course, all that is about the presentation / symptoms, and leaves open the question of underlying cause(s).
That’s very helpful, thanks. I think the key here is this part:
The short-term predictor is a learning algorithm. Given infinite time, we normally expect learning algorithms to settle into some steady-state configuration where they no longer update. We can think of this configuration as “what we are training it to do”. So, in this toy model, what are we training the short term predictor to do?
I’m trying to talk about what will happen given infinite time, in steady state, i.e. when it gets to a fixed-point / self-consistent solution. In steady-state / at a self-consistent fixed point, the STP output approximates the expectation of the next override (and/or the expectation of the STP output at the next change of context data). So we can call that “prediction”, in the sense that “it’s a signal which tells us that a certain thing will happen later”. But it’s not a “prediction” in the sense of “passively predicting an independent, exogenous event”.
Rather, the LTP is one component of a machine (that also includes the stomach or whatever), and we’re narrowly zooming into that one component and seeing that its outputs can be interpreted as predictions, at the fixed point.
Relatedly, in logical induction (cf. here, or Theorem 4.11.2 in the original paper), they formulate a seemingly-paradoxical sentence “this sentence is true if you predict it to be true with probability LESS than 50%”. Does it have a fixed-point / steady-state? Yes, 50%. And that’s what it converges to.
Another pathological case would be to make everything a fixed point: in the case at hand, we could set up a LTP where the override is by definition whatever the STP is outputting at that moment. So then the LTP will stably output whatever random value the STP was spitting out when it was randomly initialized. Nevertheless, we can still call that a kind of “prediction”, I think. Like, we look at the STP output of 1.3 and say “it’s predicting that the next override will be 1.3”, and then the next override comes, and indeed it is 1.3, which validates that point of view.
Anyway, the text didn’t make any of this very clear, because I wasn’t really thinking about this aspect of it, so again I appreciate your comments.
Indeed, I’m now questioning a bit whether “prediction” is the right word here, since the word “prediction” does usually have a connotation of “passively predicting an exogenous event”, and I have previously criticized people for using the word “prediction” in weird situations where that connotation does not apply, so it might be hypocritical if I’m doing that myself. Hmm, I think the word “prediction” is still OK here, but I would want to add some clarifying text for sure.
Compare R=1 to R=2. If these are predictions, then R=1 is a higher [predicted probability that the next override is R=0] than R=2. But R=1 produces fewer enzymes, which makes the R=0 override less likely.
Think of this in terms of what the fixed point / steady-state / self-consistent solution would be.
Let’s say context is always the same, and on day 0, the STP outputs R=2, and that’s almost always too much digestive enzymes, so there’s R=0 overrides 90% of the time (and R=10 the other 10%). Over the next week, the STP weights update to make the predictions incrementally lower, and now R=1.8, and there’s R=0 overrides 85% of the time. Over the next week, the STP weights continue to update in the same direction, until now R=1.7, and there’s R=0 overrides 83% of the time, and R=10 overrides the other 17%. And now we’re at the fixed point! And indeed the STP output is now the expectation value of the next override, so we can (maybe slightly dubiously) use the word “prediction” to describe this output.
Thanks!
On the question of what BF Skinner actually believed: you obviously know more about it than I do. Feel free to suggest something I could read to better appreciate the value of Skinner and radical behaviorism (but no promises that I would actually read it anytime soon).
Anyway, I was not trying to attribute any particular position to BF Skinner or anyone else; this is not a post about the history of academic psychology :) In footnote 3, I say that, when I was picking the terminology “behaviorist”, I was trying to invoke a “caricatured” notion of behaviorism, not necessarily a historically-accurate one. Anyway, I wanted to talk about a certain thing, and needed to make up a term for it, and “behaviorist” seemed like an OK choice, and that’s all there is to it.
Thanks so much! I think your response has some mistakes, but I think you had the right big picture, especially the idea of bounding
between 0 and ¼.Here’s my attempt, inspired by reading what you wrote + your two sources (the Zuk et al. 2012 supplementary info section 1, and the book Genetics and Analysis of Quantitative Traits pp 81ff):
Assume no assortative mating and no gene-environment interaction. Normalize the variance of the trait across the population to 1 (so “fraction of variance due to X” is the same as “variance due to X”). Then we have:
where
= contribution to variance from unique environmental effects = contribution to variance from common-between-siblings environmental effects = a genetic contribution to variance, involving an interaction of additive effects at i different loci, and dominance effects at j different loci (with the convention for nonsense combinations like ).
We split the last term into:
An additive contribution
A non-additive contribution
Importantly,
is the fraction of population variation explained by a PGS, in the limit of perfect PGS measurement (infinite sample size, and measuring every last rare variant, structural variant, copy number variant, and whatever else).Next:
Now split up the additive vs non-additive parts:
Subtract to get:
Key step: the
coefficients in the sum are all between 0 and ¼. Thus:Now substitute
(from above), and we finally get the following:Reassuringly, this agrees with both of the intuitive guesses I suggested above: assuming
, then for , we get , i.e. the PGS would (in the limit of perfect PGS measurement) work almost perfectly, with no missing heritability; and for , we get , i.e. the PGS cannot possibly work at all.
FYI, two days after I published this, Beren Millidge published How can LLM RL Work Despite Information-Theoretic Inefficiency. I don’t think it contradicts anything I wrote here; rather it complements it by asking and answering different questions. Recommended!
This is good concrete brainstorming, thanks.
Suppose, on the other hand, that the classifier is itself capable and that the smallest capabilities level required to jailbreak the classifier is higher than the classifier’s.
I don’t think that’s a good assumption. My normal assumption is that if you crank a pretty basic learning algorithm against a learned classifier, you’ll find an adversarial policy in short order. For example, a paper from 2022 achieved “a >97% win rate against KataGo running at superhuman settings … [u]sing less than 14% of the compute used to train KataGo…. Our adversaries do not win by playing Go well. Instead, they trick KataGo into making serious blunders that cause it to lose the game.”
Also RL & search systems (if they get to AGI) would be able to do “real” open-ended continual learning (see here), whereas (IMO) LLMs are basically stuck with the human conceptual space. That means that an RL & search AGI can rocket way out of distribution, by figuring out new concepts and building new technologies etc.
This causes two problems: (1) the LLM would struggle to judge whether the AGI’s behavior is good or bad, and (2) the LLMs would struggle to judge whether the AGI’s thoughts are good or bad. The latter is important because of the generic issue that (IMO) “Behaviorist” RL reward functions lead to scheming. Judging inscrutable thoughts is hard under the best of circumstances, but when the thoughts might include just-invented new conceptual spaces unlike any that have been thought by anyone before, it’s even harder.
Even in the case of just judging behavior, it’s maybe a bit like taking an LLM trained on pre-1930 text and asking it intricate questions about the ethics of social media meme accounts and so on. Maybe it could do OK if there were an honest broker filling the LLM’s context window with tons of relevant information about what all these different words mean and what the heck is that glowing screen full of rectangles. But there is no honest broker; the AGI itself has a different agenda.
My point for this post is that we need a plan for RL & search AGI, and that we don’t have one. A plan might involve a certain starting training environment in addition to a certain reward function. Sure.
If you’re suggesting that, as long as the training environment is good (e.g. a loving family with adequate socializing), then we can use some straightforward reward function (e.g. “reward when the supervisor presses the ‘approve’ button”), and the AI will still wind up being genuinely nice, then I strongly disagree. E.g. plenty of human sociopaths grow up in loving families; a paperclip maximizer would maximize paperclips regardless of its childhood environment; and when people raise undomesticated animals as beloved pets, they often wind up getting mauled.
So I think the reward function has to be a central part of the plan, and I tend to emphasize that. But yeah, whenever I say “What reward function would lead to a nice AGI?”, that’s shorthand for (e.g. here) “What reward function (along with training environment and other design choices) would lead to a nice AGI?”.
See Intro series §12.5 for further discussion.
people, unlike LLMs, don’t tend to imitate their enemies. I wonder how it interacts with tropes like The Chain of Harm and bullying traditions
This is kinda off-topic, but I’ll respond anyway.
I think the motivation to act aggressively (in certain contexts) is innate in humans, just as it’s innate in probably all complex animals. So we don’t need to explain from scratch how people develop a motivation to bully each other. The motivation is already there, and might or might not come out depending on their personality, their mood, their relationship to the potential target, and lots of other mitigating and aggravating factors.
By contrast, we do need to explain how someone might develop a motivation to, I dunno, get a cupcake tattoo. There’s no innate drive that makes it feel satisfying to get a cupcake tattoo. Instead, the explanation would probably be that a cupcake tattoo is the kind of thing that this person’s idols (real or imagined) would think is cool.
(This comment is kinda poorly thought through, sorry, but it’s the best I can do right now and I don’t want to leave you hanging even longer.)
Hmmm. It’s possible that you’re onto something important, but I’m not convinced yet. :)
(It could be 0, but only if I incorrectly start to produce more enzymes in the meantime, the STP had been incorrectly outputting a high number.)
You’re brushing this aside, but to me it’s load-bearing. The R=1 output leads to more enzymes than R=0, and maybe even that small amount will wind up being too much. Probably not, but you only need that to happen 10% of the time.
(Another degree of freedom is, instead of R being the “rate of digestive enzyme production”, it could be a more abstract “rate setting” that is monotonically but nonlinearly related to the literal number of enzyme molecules produced per second.)
(Periodic reminder that I don’t actually know anything about digestive enzymes, this is still just a toy example.)
…It might be nice to have easier-to-visualize metaphor. Let’s try!
Imagine, on a dare, you’re driving pretty fast, blindfolded, on a curvy go-kart track. You’ve driven on this track before, so you kinda remember the way it curves, but not very well, and you’re also pretty hazy on where exactly you are relative to the track.
Sometimes you hit the left side of the track, and then you scream and pull the steering wheel hard to the right (“+10 override”). Sometimes you hit the right side of the track, and then you scream and pull the steering wheel hard to the left (“–10 override”).
Otherwise, I guess my scheme would be: you’re trying to keep track of which side you’re likely to hit next, and the more you think it’s probably gonna be the right side that you hit next, the more you’re gonna turn the steering wheel left, and vice-versa.
So for example, at time t=0, you’re steering straight, and you think “I’m probably going to hit the left side next”, so you start turning the steering wheel more and more to the right, at 5°/second. Then at t=2, the steering wheel is 10° right of neutral, and now you think “I’m 50-50 on which side I’m going to hit next”, so you keep your steering wheel fixed at 10° right of neutral. Then at t=7, it’s been long enough that you have a sense that you must be getting close to a bend where the track veers left, so you think “I’m almost definitely going to hit the right side next”, and start turning the steering wheel 10°/second to the left, until the steering wheel is neutral at t=8, then 10° left of neutral at t=9, and so on until you feel an increasing chance that you’ll overcorrect and hit the left side. Etc.
See what I mean?
Anyway, when I read your comment about a scenario where (based on past experience) you’re expecting low need for digestive enzymes for a couple hour, but high need after that, I was sorta visualizing this go-kart track, following the curvy statistical expectation of how much digestive enzymes you’ll need when, and trying not to bump up against either the too-high or the too-low side. It’s important that there’s always a chance of hitting either side of the track.
Does that help?
(PS: in my first draft of this comment, I was gonna suggest that the probability of hitting the left versus right side of the track should control the steering wheel position, not the steering wheel rotational velocity. But that wouldn’t work in a circular track—it would never stop hitting the outside, I think. Seems vaguely related to P versus I feedback control, maybe? But not exactly … it’s not a traditional control loop because it involves a forward-looking prediction, I think.)
In the OP, I contrasted the “comparatively-less-pessimistic group (say, P(doom)…in the 5%–50% range”, with the “even more pessimistic group” that includes me. I think the style of argument you mention (“strong economic and military incentives for making minds that are highly agent-y…”) is a good argument for being in the former group, but it’s kinda hard to justify P(doom)>>50% that way.
For example, an optimist could respond to the “incentives” argument by saying “well humans can be pretty agent-y, and humans can accomplish ambitious projects, but yet humans also care about our friends and follow local norms and customs etc. So it’s at least possible that AIs could be like that!” And then the pessimist could respond “yeah but the more ruthless AIs will outcompete the nice ones”, and then the optimist could respond “well the vast majority of the AIs will be nice because competent companies and militaries won’t choose to run AIs that are eager to wipe out humanity, so the few omnicidal AIs can be stopped”, and then maybe the pessimist would pivot to offense-defense balance or whatever, and on we go down the argument tree. Anyway, there are some people who get to P(doom)>>50% via these kinds of arguments, but I think it’s more common that these arguments only get people into the P(doom)≈5–50% range.
And then I’m in a different cluster that includes Eliezer & Nate, and says that “P(doom)>>50% because egregious omnicidal misalignment is what’s definitely going to happen in the absence of some unlikely technical breakthrough”. That’s what I was defending in this post.
For many purposes, the “comparatively-less-pessimistic group” and the “even more pessimistic group” are on the same team, e.g. against Marc Andreesen. But in other contexts, the two groups are on opposite sides, so it’s very worthwhile to try to hash out which group is right.
Yeah, my thoughts exactly (if I understand you correctly).
I mentioned in the OP that the NVIDIA paper (Liu et al. “ProRL”) says “RL can indeed discover genuinely new solution pathways entirely absent in base models, when given sufficient training time and applied to novel reasoning tasks,” but then added that I didn’t the paper had proved it. I didn’t explain in the OP why I was skeptical.
…But what I was thinking was: if solving the problem requires doing the right step 20 times in a row, and the base model has a 10% chance of taking the right step each time, then the base model will never succeed, at least not in the number of attempts that they could afford to try. But then if RLVR gets it from 10% to 95%, it will succeed a lot. But upping the probability from 10% to 95% is not what one would reasonably call a “genuinely new solution pathway entirely absent in the base model”.
(Warning: I skimmed the paper and might be misunderstanding how they were justifying that claim.)
Hmm, I reworded the section heading
FROM “Theoretically, each GPU-hour spent on RL should have orders of magnitude less contribution to LLM capabilities than a GPU-hour spent on imitative learning”
TO “Theoretically, each GPU-hour spent on RL conveys orders of magnitude less information content than a GPU-hour spent on imitative learning”
Sorry about that.
The RLVR is a small change to the model in the grand scheme of things—that’s my point—but of course it’s a small change that makes the model much better at things that the companies care about. If there was a way to get that same small change via imitation learning, then that would require much less compute, and companies would definitely want to do that. But no such alternative is known. …Well, oh actually, there is a way to do that: you can find some other model that can already do good inference-time reasoning and then distill it (via imitation learning). And companies do exactly that whenever they can. But you can’t push the SOTA that way. So they pay the cost.
Yeah I think that’s overall a very reasonable stance on LLM alignment.
I forget. Definitely not a reliable source. In fact, I’ll edit the guess to 20% now. If anyone knows more, please share.
In the big picture, I basically agree with all that. But I’ll nitpick a bit anyway :)
Whether the model knows which pastry Baker Brun is known for would seem to me to be of ~0 relevance to alignment. Its goals, reasoning abilities and in-context learning abilities strike me as where ~100% of the meat is.
I agree that “knowledge” of Baker Brun in particular is not related to alignment. But we shouldn’t generalize from that example to saying that pretraining is irrelevant to goals. (Maybe you didn’t mean to insinuate that anyway?)
For example, a base model may well autocomplete “I’m cold” to “I’m cold, so I’m gonna put on my coat now!”, which is pretty goal-like. Granted, it’s still not a true goal yet, but once we bring in tool use, those same autocomplete expectations can turn into bona fide goal-seeking actions. And yet they’re still derived from pretraining, and still reflective of the human distribution.
I don’t think we have any reasonable bound on how quickly SFT+RLHF’s niceness gets dilluted away by RL
I don’t know how to bound it from first principles, but at least we have some empirical data by now.
Risks depend on both capabilities (could it do Bad Thing X if it wanted to?) and alignment (does it want to?).
LLM alignment: The §3.3 discussion is my take on that, and it hasn’t changed for a long time (e.g. compare with Foom & Doom §2.3 from June 2025, and the non-RLVR part of that is in turn parroting §4.2 of this post I wrote 2023 …).
LLM capabilities: I didn’t discuss this in the OP, but my opinion is still that there’s a certain kind of “figuring things out” that humans can do (especially over extended periods of time), but that LLMs can’t, not now and not ever. (I’m stating an opinion without defending it.) But it’s tricky for me to translate that hypothesized limitation into concrete predictions of what future LLMs will or won’t be able to do in the real world. Could future LLMs wipe out humans, invent science and tech centuries beyond our wildest imagination, and colonize the galaxy? I’m confident in “no”. Could LLMs wipe out humans, leaving aside the question of whether they’d be able to survive on their own afterwards? I still lean “no”, but less confident. I’m certainly happy for there to be people working on LLM x-risk, and indeed I think there should be way more work going into that. But I also think the non-LLM thing that I’m working on (“brain-like-AGI safety”) is an EVEN scarier, more neglected, and more likely x-risk on the horizon.
I edited a sentence in OP: It used to say “…Whereas Rohin is saying: LLM capabilities mostly come from imitative learning, therefore CoTs are legible, and this will not change too soon”, but now it says “…this will not change too soon, absent some important future change in LLM training approach.” I agree that this is an important caveat, thanks.
I have no opinion about whether there will be important future changes in LLM training approach. What you said sounds like a plausible consideration, sure, but I dunno.
“CoT monitorability is a fragile opportunity” seems like a fine framing to me. I mean, we can pessimistically emphasize how CoTs are not a certain panacea for safety, or alternatively we can optimistically emphasize how CoTs are not always completely useless for safety. But that’s just a vibes disagreement, because both are true.
Update: I wrote a post which expands on this comment: LLMs are (still) mostly powered by imitative learning, not RL.
Thanks!
I think this summary relies on us pretending/assuming/modeling “yes, we get an override even if we’re currently predicting the right thing”. Otherwise, the next override will always be the opposite of what we predict.
Is that assuming that the signals are binary rather than continuous? Let’s say the override ground truth says “produce digestive enzymes at rate R”, where R can vary between 0 and 10. Or to simplify, let’s say the overrides are always saying R=0 or R=10, whereas the predictions can be any real number R. Then it’s never going to predict R=10.0000 if there’s any uncertainty whatsoever in the next override, but it might predict 9.9 if the next override is 100× more likely to be R=10 than R=0. (“…if there’s randomness that makes the next override unpredictable, we are training it to predict the expectation value of the next override.”)
So if you’re expecting to eat with very high confidence, you’re nevertheless going to produce not quite enough digestive enzymes, and the last bit of digestive enzymes will come in when food actually enters your mouth.
(Similar example: my heart will race in anticipation of seeing a scary thing, and if I’m very confident that the scary thing will show up then my heart will race quite a lot, but even so my heart will race still more when the scary thing is actually there in front of me.)
I’m happy to edit the text to make this clearer (sorry), but first I want to check that I’m actually understanding and responding to the point you’re bringing up.
Not in cultures like ours with widespread photography, mirrors, and art. However, in the absence of those things, I think sometimes yes, see §7.4.3.1 here.
I think we like looking at stick figure drawings for a similar reason as we like hearing gossip. It’s not that it’s directly triggering brainstem sensory heuristics, but rather that it’s making us think about other people, which can be its own reward.