kaarelh AT gmail DOT com
Kaarel
Based on those posts, if RLVR is 50% of the training compute, then it might only be 0.01% of the information content imparted by training. And unless the learning rates are drastically different, it would also be 0.01% of changes to the weights.
I think “it would also be 0.01% of changes to the weights” is not in fact entailed? There is a zeroth order baseline that it’s more like 1:1, because when you get a successful transcript on a problem in RL, roughly speaking you just as-if-pretrain on that transcript (like, you get a gradient contribution from asking for an increase in the probability of each token of that transcript). That said, I think there are important corrections away from this.
edited to add: oops ok i see i was probably misunderstanding you: i was talking about the per token learn rate in each case; you were talking about the per RL transcript learn rate vs the per pretrain token learn rate. i would guess that my normalization is somewhat more natural as a zeroth order baseline, but not sure and mostly nvm the point above then, sorry. see item 3 of the list below
I will now list some considerations on the general topic of RL vs pretrain thinkoomph contributions off the top of my head, without putting in the work to put these together into some coherent story atm:
1. A model-generated transcript that solves a problem will be close to what the model already does naturally, so there is less for the model to learn from it — like, a human’s transcript on the same problem would exhibit “tricks” that are more novel to the model.
2. otoh: A model-generated transcript will be natural for the model, so the tricks in it are easier for the model to learn — like, a human-generated transcript for the same problem would be more alien for the model; each trick exhibited in it is harder for the model to learn/use.
To illustrate the above two points with a hyperbolic example: it’s somewhat like a teacher guiding your thinking a bit when you’re trying to solve a problem vs telling you a solution in chinese and trying to get you to repeat it back.
3. Fable’s guesses are that for frontier models, pretrain token vs RL token weight delta sizes are 3x in favor or pretrain, whereas pretrain token vs RL transcript weight delta sizes are 300x in favor of RL, with wide uncertainties.
4. Under “RL”, I think labs have in part been doing some things other than vanilla RL, and I think this could easily contribute a >10x factor to thinkoomph calculations. Some examples in this category (various of these are kinda iterated amplification and distillation):
4.1. Have a model try to solve a problem. Have a human prompt the model with corrections and other hints until it solves the problem correctly. Then train on the transcript with these hints removed. Or maybe train on the transcript with these hints, idk. (I think labs have effectively (indirectly) been employing at least on the order of 100k graduate-student-likes to provide training feedback, including this sort of thing as a significant part.)
4.2. Have the model distill a long solution path to a streamlined one that eg cuts out bad branches, then train on that streamlined one.
4.3. Have the model spend a lot of time solving a problem, then write hints, then prompt a model to solve the problem with these hints, then train on the transcript with the hints removed.
4.4. The previous item but with math textbooks or paper fragments etc in context.
4.5. Give the model the/an answer to the problem, ask it to write a way to get to that answer, then train on this rationalization.
fwiw my own high-level guess is that the understanding/skills of current models are in some sense (that is hard to make precise) more from humans than originally created, but my guess is that at this point it’s significantly less extreme than mt everest vs an elephant
(some messages i sent to a friend at coefficient giving 6 months ago, with a few edits)
maybe openphil global health and wellbeing should fund some cryonics evangelists and scalersassuming developing advanced AI goes well for humanity (which openphil effectively seems to have a decently high probability on) or we ban AI and also don’t destroy humanity/capitalism/institutions, you are likely buying > billions of QALYs for each actual current human who would have died that gets cryopreserved. (edit: hmm i guess this is assuming recovery from the outputs of a human is fucked. if you can recover someone from their writing as well then maybe there’s not much to gain from being cryofrozen. not sure what i think about this.) i think currently the usual price per person is like 100k usd but i think it could be like 10k with worse methods that are probably just fine, and at massive scale it could be like 1k — if very widespread, cryonics could plausibly even be cheaper than other usual things people do with a body
i mean it’s unclear if these cost numbers are really directly relevant in the naive way. like, you wouldn’t compare these to bednet usd/QALY numbers, because openphil wouldn’t be paying for these individuals getting cryonics probably. tho idk, i guess openphil could be initially paying for some fixed costs of setting things up which could be well estimated per person by these numbers maybe. well, openphil could just also be paying for individual people lol. anyway, my guess is that the correct calculus will make the bet look much better than just paying for each cryopreservation
For thinking about the AI situation, it helps to be familiar with a bunch of numbers. [1] Here’s a basic browser game (created by me and Claude) for practicing this: https://kaarelh.github.io/fermi/
some ways you can help improve the game:
suggest more questions
tell me if you think a question should be removed/improved/corrected
suggest changes to the scoring system or any other aspect of the game
- ↩︎
things such as AI time horizons, training costs and operation counts, energy use, various doubling times, etc
I’d suggest that you consider adding at least one case for faster takeoff to your reading list. Some options:
Another case that I unfortunately don’t know of a good writeup of is:
I think that AI 2027 thinks this: supposing we get to SAR and stop scaling compute, it would take the human research community median
years to get to ASI (like, with software alone); they conclude median takeoff time >1 year by a calculation from this param. If one takes the human time to instead be <1 year or <10 years and propagates that through their model, my guess is that one will get SAR to ASI takeoff in at most a few months, but I haven’t thought this through carefully.
- ↩︎
though I wish they focused more on [the time to get from AIs being better than top researchers at most research projects to nanobots], and less on econ params
(sorry, haven’t read most of the stuff you link, but thought this would be positive EV to say anyway)
Rather, we need to argue that A has higher “expected value” broadly speaking, meaning: In some sense we “expect” that, if we were idealized agents who could aggregate all of A’s and B’s possible consequences into literal EVs, then we’d say A has higher EV.
I don’t think we should accept this meaning of the EV of an action / this operationalization of what it is to do EV-thinking. it’s like how when I calculate 2+2, I’m not calculating what I would conclude if I were to try to figure out what 2+2 is, I’m just calculating what 2+2 is! i think you’re confusing truth and provability
like, when i make probability or expected value claims in general, from the outside one could look at me and say i’m just playing some game involving relating numbers to other numbers and betting attitudes and whatever, and maybe explain why i’m playing this game via eg jeffrey-bolker. but if from inside the system, i were to view it as legitimate to translate probability claims into some concrete claims about what sort of game i’m playing, then that would introduce all sorts of crazy stuff such as thinking that if i were to bet at 0.5 on P then that would make the probability of P 0.5. this is crazy — probabilities are supposed to be objective things from the inside, not things that can be changed by your attitudes.
(in particular, there is an important sense in which moral antirealism is false, just like truth being provability is false.)
see yudkowsky’s metaethics sequence for a more detailed version of this argument. i also recommend the book gödel without tears in case you’re not already familiar with the incompleteness phenomenon (including löb’s thm)
(i guess one could describe this as me rejecting your P1, but it feels more like i think you are saying stuff in a confused frame)
setting aside the above objection to this genre of imo confused antirealism, the following example still seems extremely bad for the specific version of the view behind your premises, though maybe i’m misunderstanding the view:
Let option
be a certainty of getting utils. Let option be utils, where is the number formed by decimal digits and of .Note that you should pick
, because its expected utils (assuming is normal and it is reasonable from your bounded perspective to have a uniform distribution for those specific digits) are .However, you should have
that an ideal version of you (who knows the digits of , and again assuming normality and a uniform distribution) would tell you to pick , because this is what you should do in case the logical variable turns out to not be .So, it seems that your view says it’s not justified to pick
over , because you would expect ideal advice to not be to pick , and to pick instead.But this seems very silly.
(You can try to fix this by speaking of your expected value of your ideal guy’s expected value of the options, not of what you expect the guy to decide, but at that point maybe it’s clear that you’re allowed to just be doing stuff with the expected value of the options directly?)
This is a special case of a general principle: the only way to know that you don’t secretly want X, desire X, or value X, is if you would be able to admit to yourself that you do want/desire/value X, in worlds where you really do. If the possibility feels too painful to face, then you can’t rule the possibility out.
i think you have a silly and wrong view of values where there is some pre-thought pre-determined secret given wanting/desiring/valuing/[utility function] that can somehow be “directly accessed”. really, your values are thoughtfully given — or more precisely your actions are properly guided by a process involving rich ethical thought — and it is in particular just completely fine for you to end up not wanting X (in every sense relevant to action) because you are blocked from this by it being really painful to contemplate the possibility that you’d want X. there’s no “you really do want X” that is somehow mysteriously already given before thinking about a question. i guess you could try to put yourself in some sort of specific state and try to access some sort of feeling there, but i see no reason to attribute this profound special significance to that
imagine a guy whose telopheme says (among other things) “you do not want X” and a guy whose supposed-telopheme says (among other things) “you want X” but then there is a totally-different-mechanism-bro on top of the telopheme that adds “do not” to this sentence, with this happening before this sentence in the supposed-telopheme starts affecting the rest of this guy’s thinking. i think the second guy is just completely fine compared to the first guy? there’s nothing special about the supposed-telopheme, no reason for special interest in it. “look deeply at your TRUE feelings for the truth bro, listen to your HEART’S DESIRE” is just some value-flattening woo woo nonsense :P [1]
In order to know your own wants, desires, or values, you need to be ready to face the possibility that they’re not the wants, desires, or values you thought you had, wished you had, or were told you were supposed to have.
of course what values you wish to have affects (and should affect) what values you have! it’s also fine and proper for what values you’re supposed to have for eg your family or friends or community to affect what values you have! again, there is no secret thing there already given before thinking (including even communal thinking with other people) gets started, and these are fine forms of ethical thought
(here’s another comment where i make the same point)
however there is some sort of “being honest with yourself and contemplating options open-mindedly is pro tanto good” here that i agree with, but as basically an ethical principle, not a meta-ethical fact
a postscript for fun: i personally find it a fairly simple mental exercise to feel sexual attraction to like basically anything, like including mathematical objects (stellated dodecahedron, anyone? alexander horned sphere? there’s some fine shape rotation exercises here), scientific theories, historical events, cultures, etc.. i’d guess that most smart people could do this if they had seriously practiced being thoughtful about their thinking when they were still able to learn. i’m almost completely uninterested in what some primordial default setting for sexual attraction is, and i don’t see any reason to revert to it. i don’t see any reason to take whatever the current setting happens to be very seriously either. i consider this a sort of silly barbaric thing to be overwritten with little concern, not a deep ethical wellspring. (however i also don’t consider it wrong if someone, as a personal ethical principle, takes this to be more of a protected source of ethical guidance. (however i do consider it a serious ethical mistake to put this on a pedestal, even only in one’s personal ethical thought.))
- ↩︎
also i don’t think valuing works at all like this telopheme picture (or at least any non-galaxybrained implementation of it) anyway — i just think of it as a simple toy model in which the point i want to make can already be made
- ↩︎
variant continuation: “Treatment is simple. Consult HCH. It should solve these problems for you.” Researcher bursts into tears. Says, “But doctor, we are HCH.”
remark: it’s wild how most proposed alignment schemes are “make a guy that solves alignment for you” (i think this is true even of many schemes which try to be principled). it’d be interesting to better understand the extent to which sth like this is necessary. [1]
oh right, because a computable bound on the average proof length would imply a computable worst case bound as well (since there are only
statements of length ), good point! I guess two remaining directions here are:can we say some more stuff about what the distribution of proof lengths is like?
is there an interesting scaling law for the statements a reasonable mathematical community (such as the human one) actually proves? (I think sth like this is what I’m most interested in here)
I generally recommend looking first at the presentation slides and then under “uncategorized AI safety”. Specifically, I suggest Model-wise thinking.
notes!
I’ve just posted a repo with a bunch of my notes from 2023–2026, mostly on topics with relevance to AI alignment; see here for more meta information. An assortment of 101 items from the vault:
AI safety presentation slides (or really self-contained slideuments):
Variants of the alignment problem (at the AFFINE Superintelligence Alignment Seminar)
Verification-based alignment schemes (at the AFFINE Superintelligence Alignment Seminar)
Model-wise thinking (at the AFFINE Superintelligence Alignment Seminar)
Impact cases for guardrails/monitoring/verification (at Mila)
on understanding:
valuing, ethics, metaethics:
in one sense, human values are very complicated; in another sense, human values are very simple
an illustration of the role of understanding-machinery in value development
trajectories of moral-reflective flight — an alternative to reflective equilibrium
some thoughts on consequentialist ethics in large worlds and an ethical problem relating to existence
ideal induction math:
on ideal inductions and attempts to use them to handle the AGI problem in principle:
philosophy of language, philosophy of science, epistemology
interpretability:
metaphysics:
on human futures, facing AGI:
uncategorized AI safety:
varia:
Unfortunately, there exist unkind people. We can select for people with eg long prosocial careers, high introspection, positive interviews with close friends / relatives, cognitomotor symptoms of empathy etc
I think a central property you don’t mention explicitly is [having kept promises in the past, especially in cases where these required doing difficult costly things and where the person thought they would not be rewarded in the future for having kept the promise]; also, honesty. I’d guess that an important class of potential successes with this kind of scheme, in fact maybe most of the successes (but idk), involve the fooming mind [keeping a promise]/[maintaining a commitment]. I think that maintaining some kind of kindness without a specific commitment to helping existing humans in some way can easily “misgeneralize” to eg some sort of utilitarianism, and nearly every kind of utilitarianism endorses the atoms and negentropy of all existing people being used for something else, or just more generally misgeneralize to caring about new people you can create and various other possible beings and activities over existing people.
also copying a note i wrote for myself on this topic: ”
some ideas for safe self-improvement
Most people currently thinking about AI alignment seem to hope that there is some sort of “formula” for safely/[value/character-preservingly]/whatever becoming more capable (and for alignment more broadly). I doubt there is some such formula to be found. Instead, I think that as one becomes more capable, one should keep thinking carefully about how to become more capable, and that there isn’t some “formula” for how to do this thinking. I think there is very much to be understood about how to become more capable “safely”. This note presents some basic ideas for that.
Why care about how to self-improve safely?
Here are some more concrete reasons to be interested in ideas for safe self-development:
You might hope to somehow make an AI at
human intelligence that kinda “properly cares about humanity” for its intelligence level, to let it self-improve until it can take over the world, and to have it “preserve” this property of properly caring about humanity throughout this self-improvement well enough that it then ends the present period of (imo) high existential risk from AI. We could imagine the ideas here being developed into an initial guidebook for such an AI. This would especially matter for the period during which the self-improving AI is still kinda dumb so it couldn’t yet write a better guidebook itself. Once it becomes kind of smart, one might hope that it would do something like continually improving this guidebook as it proceeds. I think this “plan” is crazy and shouldn’t be attempted, but many people seem to think that it would be just fine to let 2025 Claude foom or whatever, so maybe you dear reader are interested in this.mind upload case. both the step from human to mind upload and also later steps
humanity fooming together
Some ideas and questions for self-improving safely
maybe most importantly, you have to actually think about whether some idea for self-improving is fine. you shouldn’t just be doing stuff carelessly. you have to think about how to do this thinking. you will already be doing this by default when you’re doing “object-level” thinking, but i wanted to make this explicit. like, you have to be developing theory/understanding that helps you with these questions. you have to be writing ever better versions of something like this list for yourself
for example, as a mind upload, before making a bunch of clones of yourself and sharing power with them, you should try to think through analogous things you are familiar with to understand the sorts of issues you might run into (eg what are the issues with democracies? how is a society disanalogous to an individual in general)
making a new mind from scratch is an extremely extremely scary/stupid way to “self-improve”. (haha @ humanity.) you should probably basically only be doing stuff that looks much more like becoming smarter yourself — new versions should mostly have the same structure(s) as previous versions. like, you should do stuff much more like adding or switching out small parts one by one
you try self-modifications. you leave a previous version of yourself behind to analyze post-modification you, that is supposed to be able to roll back changes. i mean this might just look like terminating the version with a certain change
this works much better if you remain “honest/transparent about what you’re like to the evaluator” along the branches considered. so you should try to do that
my guess is that one should specify some initial such structure but this needs to be arbitrarily editable by some version of the guy, because otherwise it will eventually be a stupid broken formality
there could be monitors of various levels doing various things. eg you can have a more ancient version communicating its judgments to a more intermediate version who still has the ability to reroll its future
you can implement a voting system
maybe one could try to make you believe in reward/punishment after all is said and done. like an islamic suicide bomber?
basically, the mind will kinda have to believe a falsehood
but i mean this seems possible in humans. individual humans do great science and philosophy and still believe this often. humanity “believed” this for a long time
i mean maybe there’s some way to make it kinda-not-a-falsehood?
how do you maintain a belief in god over very much thinking / capability gain (and the thing being basically false)
could look at the literature on this
how do you stay committed to your partner? i mean: how do you stay in love, how do you stay friends, how do you keep wanting to do lots of stuff together?
could look at the literature on this or just say some obvious stuff here. and then see to what extent that generalizes to the various cases of interest
sth making it a good example is that it involves a lot of growth
the following is an important difference though: in this case you’re growing alone, not with what you’re supposed to stay committed to
another important difference: you will need to keep caring about entities you couldn’t really work with, couldn’t really be intellectual partners with
how does a state stay democratic?
godspeed friend ”
it’s much closer to meaning “roman” than “walnut folk” is! romania = roman-ia = wallach-ia. it does also literally mean an eastern romance speaker, which is a subset of latin speakers, and it got that meaning from just meaning romance/latin speaker earlier
almost every proposal anyone has ever made about what a good future should maximize turns out to be a different mathematical operation performed on this one field
I think this is completely false, or at least completely false if it intends to describe conceptions of good futures in general, though maybe technically partly saved by the specific word “maximize” because maybe that word would only be used by a very specific kind of guy. E.g. I think conceptions of utopia associated with the following types of ethical views will minimally be very contrived to view in this spacetime integral way: deontology, virtue ethics, liberalism, traditionalism, preference utilitarianism[1], any kind of utilitarianism that cares about structures stretching across time (eg there being played out life narratives), views caring about the beauty of large spacetime structures, thinking of a well-lived life in terms of ongoingly chosen projects, thinking of things in terms of living a worthwhile life, thinking of human society as a being that is supposed to live a long worthwhile life, thinking of stuff in terms of good ongoing development of beings (such as humans and humanity), thinking of a good future in terms of god, any view which would be contrived to think of in terms of it seeking to create a certain kind of spacetime block, any view that rejects the claim that there should be an era of thinking about stuff followed by an era of implementing stuff (as opposed to thoughtful ethical life just continuing), etc
I think a strict version of this view where you’re literally applying some sort of functional to a fun field is even false/contrived of almost all forms of welfare utilitarianism, because even those care about experiences of macroscopic beings (like, my happiness is not well-thought-of as an aggregate of happinesses of my quantum fields (or even my atoms or cells or whatever)), and usually in principle arbitrarily large ones (eg you could have a galaxy-sized happy being whose happiness is not well-thought-of as an aggregate of the happinesses of its components).- ^
at least the version that isn’t about there being many preference-seems-satisfied mental events, but about preferences actually getting satisfied
- ^

some problems i like in alignment