I’m an AGI safety / AI alignment researcher in Boston with a particular focus on brain algorithms. Research Fellow at Astera. I’m also at: Substack, X/Twitter, Bluesky, RSS, email, and more at this link. See https://sjbyrnes.com/agi.html for a summary of my research and sorted list of writing. Physicist by training. Leave me anonymous feedback here.
Steven Byrnes
I don’t know much about GFlowNets. Anyway, that paper seems very specific to LLMs, which are off-topic for this post (see Q1 & Q7), sorry.
One line of thought I have here is, there are lots of things such a human or AGI could disagree or talk about or have an interest in, how does it pick which one? I think for the human it probably comes down to some kind of subconscious status calculation, but in either case, how does the AGI do it if it doesn’t have its own status motivations or other long-term goals?
I think we do want the AGI to have long-term goals, just not to have exclusively long-term goals, at least in the §6.2.1 approach. See my old 2021 post Consequentialism & corrigibility. Again, if you or me is the model to be inspired by, then I assume we both care about the future of life being great (long-term goal), but we both also enjoy figuring things out in the here and now, and we both also have principles that we take pride in. And for my part, I wouldn’t want to be benevolent dictator of the universe even if I could, that’s way too much responsibility, sounds terrifying.
To be clear, I’m generally expecting the process to be kinda messy, where it’s kinda hard to reason about where the AGI winds up. From my perspective, Step 1 is to have any plan at all that could plausibly work, and then we can move on to making it easier to test and de-risk the plan in advance, to the extent possible. As mentioned at the bottom, I’m still hard at work trying to get more clarity, to the extent possible, in order to make the process of AGI motivation development more predictable and legible and less messy.
I don’t think that theory [social status] explains quite as much as those people think it does
Do you have a link/explanation for this? I think this may be fairly cruxy, because I’m guessing your intuitions for truth-seeking disagreeable nerd AGI are substantially based on truth-seeking disagreeable nerd humans, so it matters what those humans’ real motivations are.
Not all in one place, but here’s some pointers (and feel free to ask follow-ups).
Let’s start with the Hansonian notion of strategic self-deception (as distinct from plain old motivated reasoning which is real and important), which I assume he got from Robert Trivers. I broadly reject that notion. E.g. here’s a footnote in my post about laughter:
Since the works of Robin Hanson are popular on this forum, I will say a bit more about where I differ from Elephant In The Brain. My biggest complaint is the part where they say:
As we mentioned earlier, people are profoundly ignorant about laughter’s meaning and purpose (at least in our default state, before learning the science). But where does this ignorance come from? Why does introspection fail us so spectacularly here?
It’s not simply because laughter is involuntary, outside our conscious control. Flinching, for example, is also involuntary, and yet we understand perfectly well why we do it: to protect ourselves from getting hit. Thus our ignorance about laughter needs further explanation.
I disagree that it “needs further explanation”. I think we start out ignorant of literally everything, until we learn it / figure it out. And I think that figuring out the evolutionary purpose of laughter is just inherently much harder than figuring out the evolutionary purpose of flinching. It’s less obvious / salient, for various reasons that I claim are pretty obvious if you think about it. I don’t think there’s any more to it than that.
I also don’t think there can be more to it than that. To explain what I mean by that, imagine if I said: “Here’s the source code for training an image-classifier ConvNet from random initialization using uncontrolled external training data. Can you please edit this source code so that the trained model winds up confused about the shape of Toyota Camry tires specifically?” The answer is: “Nope. Sorry. There is no possible edit I can make to this PyTorch source code such that that will happen.” By the same token, even if, as that book argues, there is a strong evolutionary pressure to make humans specifically confused about the evolutionary purpose of laughter, I don’t think there is any possible genetic change that would make that happen. Related discussion here.
Up a level, this strategic self-deception idea comes out of the “evolved modularity” framework in evolutionary psychology, and I reject that whole broader framework as well. See §1.1 of “My take on Jacob Cannell’s take on AGI safety” for the background, and “Learning from scratch” in the brain for why I don’t buy it (kinda related to the thing above about PyTorch).
For related reasons, I reject the idea that “status” (per se) could possibly be an innate goal. It’s just too abstract. Copying from here:
Explaining how human social instincts work is tricky mainly because of the “symbol grounding problem”. In brief, everything we know—all the interlinked concepts that constitute our understanding of the world and ourselves—is created “from scratch” in the cortex by a learning algorithm, and thus winds up in the form of a zillion unlabeled data entries like “pattern 387294 implies pattern 579823 with confidence 0.184”, or whatever. Yet certain activation states of these unlabeled entries—e.g., the activation state that encodes the fact that Jun just told me that Xiu thinks I’m cute—need to somehow trigger social instincts in the Steering Subsystem. So there must be some way that the brain can “ground” these unlabeled learned concepts.
Now, it’s not that there’s no way to solve the symbol grounding problem to make someone want social status—indeed, that obviously happens, and in my post Neuroscience of human social instincts: a sketch, I attempt to explain how. It’s that there’s no way to solve the symbol grounding problem to make someone want social status specifically. Realistically, the genetic mechanism is just not gonna be that specific. (More on which shortly.)
So here’s where we’re at so far: (1) without Trivers-style self-deception, we have a harder time explaining away introspective reports like “That’s not status-seeking because I’m not trying to impress anyone” (we can still try to explain it away, but it’s harder); and (2) the idea that “status” could have a special place as an innate end-goal is implausible anyway.
That brings us to my Social drives 2: “Approval Reward”, from norm-enforcement to status-seeking, which is a follow-up to Neuroscience of human social instincts: a sketch where I attempt to connect the neuroscience to everyday life:
That post’s §2 is the part that’s closest to status-seeking: people are motivated to have actual interactions where an actual person has positive associations with you. But even in this case, status-seeking is just one of several consequences of the same innate drive. The others are credit-seeking / blame-avoidance, and norm-following / norm-enforcement. I understand that you can try to unify these by hypothesizing that (say) credit-seeking is a means-to-an-end for achieving social status, but that’s just not how it works in my neuroscience model: credit-seeking / blame avoidance, status-seeking, and norm-enforcement all emerge in the same way, at the same level, via the same mechanism.
And then that post’s §3 gets even more distant from the conventional notion of status-seeking, by analyzing how we can feel pride in ourselves and our actions, and how this is another direct consequence of the same innate drive, and not a secret means-to-an-end to achieving social status.
And indeed, it’s easy to come up with cases where pride vs future-status come apart, and where people follow the former over the latter. E.g. people sometimes stand up for principles even if they think everyone will scorn them for it, because they’re following their own moral compass. Certainly the moral compass has something to do with what other people have said and thought over the course of the person’s life, but the relation can be quite indirect (e.g. people may care about how a cartoon character would judge their behavior), and not well-described as “trying to wind up with high social status”, consciously or unconsciously.
CoT that’s legible isn’t good evidence that LLMs are not wielding (or could not wield) novel latent abstractions/reasoning primitives/control heuristics.
I feel like you’re saying: LLMs can wield more than zero novel latent abstractions/reasoning primitives/control heuristics while CoT remains generally legible.
Whereas what I’m saying is: LLMs cannot be totally transformed by RLVR while CoT remains generally legible.
These aren’t contradictory. I think both are true.
As an example, think about humans learning things, like a teen going from her first number theory class as a teen at time 0, to deeply understanding very advanced math (e.g. the Langlands program) as an adult at time T = many years later. It’s an arduous and time-consuming process. And her notes at time T would be deeply, deeply inscrutable from the perspective of her former teen self at time 0—even the notes that are in the form of words rather than symbols.
This suggests that the delta between pre-RLVR vs post-RLVR LLMs is much less of a wrenching change than the delta between the teen at time 0 vs the now-adult mathematician at time T.
Novelty builds heavily on existing frameworks and often involves a subtle reframe, reconfiguration, or saliencing of a known/slightly modified concept in a new setting. Insight is about discovering relevance, but what’s available to be relevant is often familiar primitives.
Mathematical work is a great example here, since mathematical discoveries very often have a subtle kernel of insight, a small new idea, that reconfigures existing concepts around a problem in a significant way to illuminate something unknown. The reconfiguration is almost all in terms of familiar stuff—the moving pieces don’t change much. In the case of LLMs, that’s the content imitative learning provides.
I think instead of “novelty” here you should have said “a sufficiently small increment of novelty”.
If we instead consider mathematics as a collective human enterprise, it went from “number theory doesn’t exist at all” to the Langlands program, over the course of 200 years.
Maybe it sounds absurd for me to compare what one RLVR training can do, versus the whole edifice of ideas painstakingly built by the mathematics community over the course of 200 years. But it’s not absurd: AlphaZero really did blow past the whole edifice of ideas painstakingly built over the course of centuries by the chess and go communities in its 72-hour training runs. So the idea of building real new knowledge at a massive scale through RL is not absurd on its face. I’m just saying: RLVR-on-LLMs is not doing anything like that, at least not today. Rather, the LLM approach is to use imitative learning to suck in the whole edifice of ideas painstakingly built by humans, and then tweak it a bit at the end via RLVR.
Relatedly, when mathematicians study the recent LLM-generated math results, they’re not “learning something new” in a way that’s analogous to that teen spending years poring over her math textbooks. Rather they’re “learning something new” in a way that’s analogous to some guy telling me what his name is. If the mathematicians already have all the right background knowledge, they can quickly understand the solution within their existing conceptual frameworks. Or if they don’t already have all the right background knowledge, they can read human-created textbooks to get it. (I’m mainly thinking of the unit distance conjecture; in other examples that I looked into like the Jacobian conjecture, IIUC, the CoTs weren’t released, so mathematicians are still be a bit puzzled about how the LLMs came upon the answer.)
…LLMs too, seem to perform much of their computation through nonlinguistic representations and compress the actionable results into language, mostly because of the structural thing that language is the medium through which they maintain serial state and communicate.
This is an argument against
if you create AI capabilities via imitative learning, you get models that follow the human distribution of outputs
I don’t think it’s an argument against that, or sorry if I’m misunderstanding. You’re saying that LLMs may be computing their outputs in a different way from humans. Fine. But their outputs are still following the human distribution of outputs. “Following the human distribution of outputs” is just another way to say “low perplexity”, right?
(I will concede that the phrase “the human distribution of outputs” is a bit misleading on various other grounds, e.g. LLMs may generalize OOD in a different way from humans; not all imitative learning training tokens are created by humans; etc.)
Do you think that a human “intrinsically cares about the welfare of a photo of a human”?
Not in cultures like ours with widespread photography, mirrors, and art. However, in the absence of those things, I think sometimes yes, see §7.4.3.1 here.
I think we like looking at stick figure drawings for a similar reason as we like hearing gossip. It’s not that it’s directly triggering brainstem sensory heuristics, but rather that it’s making us think about other people, which can be its own reward.
Do you just mean that none of those other drives are necessary to make the AGI safe, or do you mean that they should all be omitted from the (first) AGI?
I meant the latter, although I don’t feel very strongly.
In fact, hmm, if we’re making an AGI singleton sovereign (either on purpose [cf. §6.2.3] or by accident), then anything we leave out of its value system will quite possibly get left out of the whole future forever (unless it defers to humans), so if the AGI feels no connection to laughter and love and so on, that’s at least a potential problem. Not sure how it balances against other things.
But yeah, if it’s just one short-lived AGI, I don’t think its subjective experience is super-high on my list of concerns, as an end in itself. I’m much more concerned about its decisions that impact the whole future. (Of course, its subjective experience is relevant to how it makes decisions.)
how do you think the AGI should feel about images of/representing humans? e.g. photos, drawings, and cartoons? Should the AGI’s innate drives trigger on them or not? Should it treat photorealistic images differently from cartoons?
I don’t want an AGI to intrinsically care about the welfare of a photo of a human. Do you? We can talk about why it might be tricky to make that happen, but the ideal goal seems clear.
(I recall that human empathy can trigger on very simple drawings, but I don’t know how much of that is directly from human innate drives and how much is generalized.)
Are you thinking of the Heider & Simmel 1944 animations by any chance?
even if the bad AGI has the option to wipe out humanity with a surprise attack with bioweapons etc, if the good AGI would be able to survive, counterattack, and destroy the bad AGI afterwards, then the bad AGI might not want to attack in the first place
A second-strike-capable “good” AGI might be an adequate solution to the problem of “bad” rogue AGI, but I think that would wind up involving the “good” AGI doing pretty scary and aggressive things: probably the “good” AGI needs to get itself running on a majority of compute on Earth, harden that compute against cyber and physical attack, make sure both those chips and those defense systems are not reliant on human infrastructure (e.g. robots defending the chips or whatever), do as much recursive self-improvement as possible, and so on. Otherwise a power-seeking misaligned AGI would be able to crush the “good” AGI along with the humans.
So I see that as a plausible solution to the “strategic” problem of §6.1, but like all the other solutions I know of, it would (1) require AGI superpowers to make it happen, and (2) have aspects that I think would strike people as quite scary and unpalatable. You’re welcome to disagree with any of that, of course.
Thanks!
I listened to a Spencer Greenberg podcast on OCPD a few weeks ago and it seems pretty consistent with (and complementary to) what you wrote.
Of course, all that is about the presentation / symptoms, and leaves open the question of underlying cause(s).
That’s very helpful, thanks. I think the key here is this part:
The short-term predictor is a learning algorithm. Given infinite time, we normally expect learning algorithms to settle into some steady-state configuration where they no longer update. We can think of this configuration as “what we are training it to do”. So, in this toy model, what are we training the short term predictor to do?
I’m trying to talk about what will happen given infinite time, in steady state, i.e. when it gets to a fixed-point / self-consistent solution. In steady-state / at a self-consistent fixed point, the STP output approximates the expectation of the next override (and/or the expectation of the STP output at the next change of context data). So we can call that “prediction”, in the sense that “it’s a signal which tells us that a certain thing will happen later”. But it’s not a “prediction” in the sense of “passively predicting an independent, exogenous event”.
Rather, the LTP is one component of a machine (that also includes the stomach or whatever), and we’re narrowly zooming into that one component and seeing that its outputs can be interpreted as predictions, at the fixed point.
Relatedly, in logical induction (cf. here, or Theorem 4.11.2 in the original paper), they formulate a seemingly-paradoxical sentence “this sentence is true if you predict it to be true with probability LESS than 50%”. Does it have a fixed-point / steady-state? Yes, 50%. And that’s what it converges to.
Another pathological case would be to make everything a fixed point: in the case at hand, we could set up a LTP where the override is by definition whatever the STP is outputting at that moment. So then the LTP will stably output whatever random value the STP was spitting out when it was randomly initialized. Nevertheless, we can still call that a kind of “prediction”, I think. Like, we look at the STP output of 1.3 and say “it’s predicting that the next override will be 1.3”, and then the next override comes, and indeed it is 1.3, which validates that point of view.
Anyway, the text didn’t make any of this very clear, because I wasn’t really thinking about this aspect of it, so again I appreciate your comments.
Indeed, I’m now questioning a bit whether “prediction” is the right word here, since the word “prediction” does usually have a connotation of “passively predicting an exogenous event”, and I have previously criticized people for using the word “prediction” in weird situations where that connotation does not apply, so it might be hypocritical if I’m doing that myself. Hmm, I think the word “prediction” is still OK here, but I would want to add some clarifying text for sure.
Compare R=1 to R=2. If these are predictions, then R=1 is a higher [predicted probability that the next override is R=0] than R=2. But R=1 produces fewer enzymes, which makes the R=0 override less likely.
Think of this in terms of what the fixed point / steady-state / self-consistent solution would be.
Let’s say context is always the same, and on day 0, the STP outputs R=2, and that’s almost always too much digestive enzymes, so there’s R=0 overrides 90% of the time (and R=10 the other 10%). Over the next week, the STP weights update to make the predictions incrementally lower, and now R=1.8, and there’s R=0 overrides 85% of the time. Over the next week, the STP weights continue to update in the same direction, until now R=1.7, and there’s R=0 overrides 83% of the time, and R=10 overrides the other 17%. And now we’re at the fixed point! And indeed the STP output is now the expectation value of the next override, so we can (maybe slightly dubiously) use the word “prediction” to describe this output.
Thanks!
On the question of what BF Skinner actually believed: you obviously know more about it than I do. Feel free to suggest something I could read to better appreciate the value of Skinner and radical behaviorism (but no promises that I would actually read it anytime soon).
Anyway, I was not trying to attribute any particular position to BF Skinner or anyone else; this is not a post about the history of academic psychology :) In footnote 3, I say that, when I was picking the terminology “behaviorist”, I was trying to invoke a “caricatured” notion of behaviorism, not necessarily a historically-accurate one. Anyway, I wanted to talk about a certain thing, and needed to make up a term for it, and “behaviorist” seemed like an OK choice, and that’s all there is to it.
Thanks so much! I think your response has some mistakes, but I think you had the right big picture, especially the idea of bounding
between 0 and ¼.Here’s my attempt, inspired by reading what you wrote + your two sources (the Zuk et al. 2012 supplementary info section 1, and the book Genetics and Analysis of Quantitative Traits pp 81ff):
Assume no assortative mating and no gene-environment interaction. Normalize the variance of the trait across the population to 1 (so “fraction of variance due to X” is the same as “variance due to X”). Then we have:
where
= contribution to variance from unique environmental effects = contribution to variance from common-between-siblings environmental effects = a genetic contribution to variance, involving an interaction of additive effects at i different loci, and dominance effects at j different loci (with the convention for nonsense combinations like ).
We split the last term into:
An additive contribution
A non-additive contribution
Importantly,
is the fraction of population variation explained by a PGS, in the limit of perfect PGS measurement (infinite sample size, and measuring every last rare variant, structural variant, copy number variant, and whatever else).Next:
Now split up the additive vs non-additive parts:
Subtract to get:
Key step: the
coefficients in the sum are all between 0 and ¼. Thus:Now substitute
(from above), and we finally get the following:Reassuringly, this agrees with both of the intuitive guesses I suggested above: assuming
, then for , we get , i.e. the PGS would (in the limit of perfect PGS measurement) work almost perfectly, with no missing heritability; and for , we get , i.e. the PGS cannot possibly work at all.
FYI, two days after I published this, Beren Millidge published How can LLM RL Work Despite Information-Theoretic Inefficiency. I don’t think it contradicts anything I wrote here; rather it complements it by asking and answering different questions. Recommended!
This is good concrete brainstorming, thanks.
Suppose, on the other hand, that the classifier is itself capable and that the smallest capabilities level required to jailbreak the classifier is higher than the classifier’s.
I don’t think that’s a good assumption. My normal assumption is that if you crank a pretty basic learning algorithm against a learned classifier, you’ll find an adversarial policy in short order. For example, a paper from 2022 achieved “a >97% win rate against KataGo running at superhuman settings … [u]sing less than 14% of the compute used to train KataGo…. Our adversaries do not win by playing Go well. Instead, they trick KataGo into making serious blunders that cause it to lose the game.”
Also RL & search systems (if they get to AGI) would be able to do “real” open-ended continual learning (see here), whereas (IMO) LLMs are basically stuck with the human conceptual space. That means that an RL & search AGI can rocket way out of distribution, by figuring out new concepts and building new technologies etc.
This causes two problems: (1) the LLM would struggle to judge whether the AGI’s behavior is good or bad, and (2) the LLMs would struggle to judge whether the AGI’s thoughts are good or bad. The latter is important because of the generic issue that (IMO) “Behaviorist” RL reward functions lead to scheming. Judging inscrutable thoughts is hard under the best of circumstances, but when the thoughts might include just-invented new conceptual spaces unlike any that have been thought by anyone before, it’s even harder.
Even in the case of just judging behavior, it’s maybe a bit like taking an LLM trained on pre-1930 text and asking it intricate questions about the ethics of social media meme accounts and so on. Maybe it could do OK if there were an honest broker filling the LLM’s context window with tons of relevant information about what all these different words mean and what the heck is that glowing screen full of rectangles. But there is no honest broker; the AGI itself has a different agenda.
My point for this post is that we need a plan for RL & search AGI, and that we don’t have one. A plan might involve a certain starting training environment in addition to a certain reward function. Sure.
If you’re suggesting that, as long as the training environment is good (e.g. a loving family with adequate socializing), then we can use some straightforward reward function (e.g. “reward when the supervisor presses the ‘approve’ button”), and the AI will still wind up being genuinely nice, then I strongly disagree. E.g. plenty of human sociopaths grow up in loving families; a paperclip maximizer would maximize paperclips regardless of its childhood environment; and when people raise undomesticated animals as beloved pets, they often wind up getting mauled.
So I think the reward function has to be a central part of the plan, and I tend to emphasize that. But yeah, whenever I say “What reward function would lead to a nice AGI?”, that’s shorthand for (e.g. here) “What reward function (along with training environment and other design choices) would lead to a nice AGI?”.
See Intro series §12.5 for further discussion.
people, unlike LLMs, don’t tend to imitate their enemies. I wonder how it interacts with tropes like The Chain of Harm and bullying traditions
This is kinda off-topic, but I’ll respond anyway.
I think the motivation to act aggressively (in certain contexts) is innate in humans, just as it’s innate in probably all complex animals. So we don’t need to explain from scratch how people develop a motivation to bully each other. The motivation is already there, and might or might not come out depending on their personality, their mood, their relationship to the potential target, and lots of other mitigating and aggravating factors.
By contrast, we do need to explain how someone might develop a motivation to, I dunno, get a cupcake tattoo. There’s no innate drive that makes it feel satisfying to get a cupcake tattoo. Instead, the explanation would probably be that a cupcake tattoo is the kind of thing that this person’s idols (real or imagined) would think is cool.
(This comment is kinda poorly thought through, sorry, but it’s the best I can do right now and I don’t want to leave you hanging even longer.)
Hmmm. It’s possible that you’re onto something important, but I’m not convinced yet. :)
(It could be 0, but only if I incorrectly start to produce more enzymes in the meantime, the STP had been incorrectly outputting a high number.)
You’re brushing this aside, but to me it’s load-bearing. The R=1 output leads to more enzymes than R=0, and maybe even that small amount will wind up being too much. Probably not, but you only need that to happen 10% of the time.
(Another degree of freedom is, instead of R being the “rate of digestive enzyme production”, it could be a more abstract “rate setting” that is monotonically but nonlinearly related to the literal number of enzyme molecules produced per second.)
(Periodic reminder that I don’t actually know anything about digestive enzymes, this is still just a toy example.)
…It might be nice to have easier-to-visualize metaphor. Let’s try!
Imagine, on a dare, you’re driving pretty fast, blindfolded, on a curvy go-kart track. You’ve driven on this track before, so you kinda remember the way it curves, but not very well, and you’re also pretty hazy on where exactly you are relative to the track.
Sometimes you hit the left side of the track, and then you scream and pull the steering wheel hard to the right (“+10 override”). Sometimes you hit the right side of the track, and then you scream and pull the steering wheel hard to the left (“–10 override”).
Otherwise, I guess my scheme would be: you’re trying to keep track of which side you’re likely to hit next, and the more you think it’s probably gonna be the right side that you hit next, the more you’re gonna turn the steering wheel left, and vice-versa.
So for example, at time t=0, you’re steering straight, and you think “I’m probably going to hit the left side next”, so you start turning the steering wheel more and more to the right, at 5°/second. Then at t=2, the steering wheel is 10° right of neutral, and now you think “I’m 50-50 on which side I’m going to hit next”, so you keep your steering wheel fixed at 10° right of neutral. Then at t=7, it’s been long enough that you have a sense that you must be getting close to a bend where the track veers left, so you think “I’m almost definitely going to hit the right side next”, and start turning the steering wheel 10°/second to the left, until the steering wheel is neutral at t=8, then 10° left of neutral at t=9, and so on until you feel an increasing chance that you’ll overcorrect and hit the left side. Etc.
See what I mean?
Anyway, when I read your comment about a scenario where (based on past experience) you’re expecting low need for digestive enzymes for a couple hour, but high need after that, I was sorta visualizing this go-kart track, following the curvy statistical expectation of how much digestive enzymes you’ll need when, and trying not to bump up against either the too-high or the too-low side. It’s important that there’s always a chance of hitting either side of the track.
Does that help?
(PS: in my first draft of this comment, I was gonna suggest that the probability of hitting the left versus right side of the track should control the steering wheel position, not the steering wheel rotational velocity. But that wouldn’t work in a circular track—it would never stop hitting the outside, I think. Seems vaguely related to P versus I feedback control, maybe? But not exactly … it’s not a traditional control loop because it involves a forward-looking prediction, I think.)
RL & search is a terrifying way to build AGI (an FAQ)
In the OP, I contrasted the “comparatively-less-pessimistic group (say, P(doom)…in the 5%–50% range”, with the “even more pessimistic group” that includes me. I think the style of argument you mention (“strong economic and military incentives for making minds that are highly agent-y…”) is a good argument for being in the former group, but it’s kinda hard to justify P(doom)>>50% that way.
For example, an optimist could respond to the “incentives” argument by saying “well humans can be pretty agent-y, and humans can accomplish ambitious projects, but yet humans also care about our friends and follow local norms and customs etc. So it’s at least possible that AIs could be like that!” And then the pessimist could respond “yeah but the more ruthless AIs will outcompete the nice ones”, and then the optimist could respond “well the vast majority of the AIs will be nice because competent companies and militaries won’t choose to run AIs that are eager to wipe out humanity, so the few omnicidal AIs can be stopped”, and then maybe the pessimist would pivot to offense-defense balance or whatever, and on we go down the argument tree. Anyway, there are some people who get to P(doom)>>50% via these kinds of arguments, but I think it’s more common that these arguments only get people into the P(doom)≈5–50% range.
And then I’m in a different cluster that includes Eliezer & Nate, and says that “P(doom)>>50% because egregious omnicidal misalignment is what’s definitely going to happen in the absence of some unlikely technical breakthrough”. That’s what I was defending in this post.
For many purposes, the “comparatively-less-pessimistic group” and the “even more pessimistic group” are on the same team, e.g. against Marc Andreesen. But in other contexts, the two groups are on opposite sides, so it’s very worthwhile to try to hash out which group is right.
Yeah, my thoughts exactly (if I understand you correctly).
I mentioned in the OP that the NVIDIA paper (Liu et al. “ProRL”) says “RL can indeed discover genuinely new solution pathways entirely absent in base models, when given sufficient training time and applied to novel reasoning tasks,” but then added that I didn’t the paper had proved it. I didn’t explain in the OP why I was skeptical.
…But what I was thinking was: if solving the problem requires doing the right step 20 times in a row, and the base model has a 10% chance of taking the right step each time, then the base model will never succeed, at least not in the number of attempts that they could afford to try. But then if RLVR gets it from 10% to 95%, it will succeed a lot. But upping the probability from 10% to 95% is not what one would reasonably call a “genuinely new solution pathway entirely absent in the base model”.
(Warning: I skimmed the paper and might be misunderstanding how they were justifying that claim.)
Hmm, I reworded the section heading
FROM “Theoretically, each GPU-hour spent on RL should have orders of magnitude less contribution to LLM capabilities than a GPU-hour spent on imitative learning”
TO “Theoretically, each GPU-hour spent on RL conveys orders of magnitude less information content than a GPU-hour spent on imitative learning”
Sorry about that.
The RLVR is a small change to the model in the grand scheme of things—that’s my point—but of course it’s a small change that makes the model much better at things that the companies care about. If there was a way to get that same small change via imitation learning, then that would require much less compute, and companies would definitely want to do that. But no such alternative is known. …Well, oh actually, there is a way to do that: you can find some other model that can already do good inference-time reasoning and then distill it (via imitation learning). And companies do exactly that whenever they can. But you can’t push the SOTA that way. So they pay the cost.
Yeah I think that’s overall a very reasonable stance on LLM alignment.
Yeah I agree that people seem to treat anthropic arguments as a special category instead of one piece of evidence among many, I was arguing about that once here.