I’m currently an independent AI Alignment researcher at Meridian Impact CIC in Cambridge, UK, formerly a staff artificial intelligence engineer and researcher working with AI and LLMs. I’ve been interested in AI alignment, safety and interpretability for the last 17 years, and have been writing about these on LessWrong for 4 years. I did research at MATS summer 2025, and PIBBSS summer 2026. I also have post-graduate experience in Theoretical Physics and an interest in Evolutionary Biology.
RogerDearnaley
This seems unsurprising to me. Alignment midtraining is useful for setting initial priors before further alignment training, particularly in ways involving a very wide distribution of situations or setting a prior for a complex, nuanced behavior, where the ability to use a lot of midtraining data drawn from a wide range is helpful.
Or, as you put it:
Much of midtraining’s effects rely heavily on seeing diverse and concrete examples of how a rule should be acted out.
But it remains the case that subsequent finetuning (or indeed RL) can shift priors sharply, so long as this shift is simple to express in the context of the world model and categorization the model has learned before the finetuning.
In your experimental scenario, 80k tokens of a data mix 98% of which does not make the distinction between the coin and charter motivations clear (so whose effect should be neutral on this issue), and 2% of which does make this clear and supports the coin motivation interpretation, containing a total of 164 examples supporting this interpretation, is sufficient to shift the 90:7 prior that alignment midtraining had established (that’s roughly a one order of magnitude prior), to 13:82 (nearly an order of magnitude the other way). Taking the log-base-2 of the ratio, this size of a shift requires only about 6–7 bits of Bayesian evidence. So 164 examples (repeated for 2–4 epochs, if I am reading your paper correctly?) gave you a behavioral shift equivalent to only 6–7 Bayesian bits of evidence. That actually seems like quite a poor sample efficiency! Would the model have been this resistant to learning this new fact without your AMT?
Those would be my first guess too. Or indeed he calls up Jensen and asks him who would be good.
There remains a possibility that Trump, perhaps accidentally, appoints someone with a functioning brain (the wording ‘AI Force’ sounds military, and if he appoints someone senior from the military/intelligence the odds of this increase), who does their research and gets scared. Trump also clearly likes it when people working for him fight each other, which can make his picks more varied.
Press reports claim Trump and his family own a lot of stock in AI companies. The question then becomes, which ones? NVIDIA seems like an obvious candidate.
There seems to have been a what might turn out to be the start of a pivot:
Trump says US will form ‘AI Force’ and appoint an artificial intelligence tsar
(Trump is also still saying full steam ahead, safety is a hoax, etc.). But this seems like an admission that this issue isn’t a nothingburger that can be handled just from the seat of Trump’s pants. Personnel is policy, so the question becomes who is appointed. Given precedent on Trump’s appointments during this administration, and under current circumstances, I’m not optimistic, but this remains to be seen.
Time is short, so obviously it’s best not to spend time reinventing things that already exist. Some suggestions of fields to look at for inspiration:
-
Biologists have done a lot of work with agentic systems and minds over the last century-plus, including building mathematical models of them. So have sociologists, and economists.
-
Alignment involves intelligent agents, morality, and incentives. Evolutionary moral psychology, the sociology of morality, and moral anthropology have done significant scientific work in this area. These are focused on systems undergoing evolutionary and/or cultural change: machine learning brings in other similar-but-not-identical processes. These also help understand that human-derived biases that distilling agency from humans into LLM base models brings along.
-
OpenAI in particular has a long history of its CEO talking about safety, then its safety-minded researchers leaving and going to found/join other companies, principally Anthropic. They’ve had a steady brain-drain of safety researchers. From what I hear, given recent events, there are now a lot of very-safety-concerned capabilities researchers at OpenAI.
FWIW, I mostly agree with OP that “effective long-horizon continual learning is likely the last remaining step to superintelligence”: there are at least a half-dozen widespread/well-known suggestions on what X might be, and for most of these I personally don’t think it’s likely to delay us long once the labs put their minds and funds to it. But this one we actually know is conceptually hard, because continual learning in human brains is a Heath-Robinson architectural contraption with at least half-a-dozen separate modules bolted together, people have been working on this for many years, and while we’ve made some progress, we’re still a long way from as-good-as-a-human.
I personally think the timeline is underestimated. This strikes me as a “me too, with bells on” piece.
Even within mathematics, what AI has been doing recently has been somewhat limited: it finds counterexamples a lot better than it finds proofs, makes interesting conjectures, or builds understanding of an area. The recent spectacular results have almost all been counterexamples to well-known long-standing conjectures, based on exhaustively applying techniques or ideas developed by others. Which is not nothing, but it’s also not all of mathematics.
People keep forgetting Amdahl’s law. If you fully automate a 1/X proportion of the work, you can speed up overall progress by a maximum of a factor of X — the part you couldn’t automate becomes the bottleneck. Analysis from Anthropic and OpenAI have both made it clear that currently the proportion of the total tasks involved in AI capabilities research they can automate is maybe up to a half, so they’re at most doubling the rate of progress. Thoe were, of course, the easy-to-automate half. AI skills are very spiky, so the easy half happens a good way before the hardest bits — likely several years before at current rates of progress, so probably less as progress accelerates.
So I think we have low-single-digit years, not single-digit months, before FOOM. Which is plenty short enough to be scary.
Another significant concern is that is that capabilities increase accelerates faster than prosaic alignment improvement, since it’s easier to automate more of, and as a result we get a serious loss-of-control before the FOOM happens.
I also distrust SamA.‘s motives, especially around the truth. But he has a motivation that I think some people forget: if many of OpenAI’s researchers think SamA. is underplaying safety, they leave, and generally go to Anthropic — as has been happening at a steady trickle (especially among his safety researchers). So one of SamA.‘s strong motivations is to be seen by his safety-concerned researchers as taking safety seriously. By all the accounts that I’ve heard from people who talk to frontier-lab researchers, a large proportion of OpenAI’s (and other frontier labs’) researchers were rightly extremely scared by recent events. The word ‘panic’ keeps getting thrown around. If SamA. just blew this off, there could well be yet another exodus of OpenAI researchers, quite possibly even a cascading one that left them seriously short on talent. So, whatever SamA. thinks in private, he has to put his taking-safety-seriously hat on. Which is a fine hat with feathers in it — he’s said all the right things periodically for years: he merely has a poor history when it comes to following that up with expensive action.
What we need is for his internal researchers to hold his feet to the fire for the expensive followup actions. Embedded external monitors from METR/etc might actually be very helpful there: they should be able into make it very hard to quietly back out of important commitments later. And SamA has already publicly committed to those.
[It is of course possible that SamA is not an idiot, does not want to die, is at least partially safety-pilled (he can certainly quote the arguments), and that he always intended to slow down once he believed systems had become actively dangerous (which we’re clearly getting very close to). But his publicly-known actions around funding for AGI safety research that Ilya wanted done, such as reneging on his 20% of compute promise, are not encouraging evidence.]
This has been discussed pretty extensively under Value Learning and Value Drift.
OP = “Original Post”
Since you’re into the math side, this might also interest you (it’s heavy on game theory):
• Binmore, K. (2005). Natural justice. Oxford University Press.
(FWIW, I think there’s a little less evidence for the Nash bargaining solution being the unique optimum than he seems to, so I’m less convinced by some of the mathematical superstructure he builds in his last few chapters 9–11 derived from that premise — but then he’s pretty clear that it’s all just a simplified toy model, so he’s not entirely convinced either. The overall thrust of his argument up to that point I basically agree with, and I think most evolutionary psychologists would agree with.)
From a paper I’m working on;
Appendix C: Introductory Reading for Evolutionary Psychology
Some useful introductory reading on Evolutionary Psychology, for anyone who might find it useful for AI Safety/Alignment (both what we’re aligning, and what we’re aligning it to), or informative for Moral Philosophy or Philosophy of Mind:
I found Richard Dawkins’ (1976) popularization The Selfish Gene very inspiring, it is very readable and covers the fundamentals such as inclusive fitness clearly. It is an early work, but newer editions such as:
• Dawkins, R. (2026). The selfish gene (50th anniversary ed.). Oxford University Press.
have been updated and include two additional chapters, including one specifically covering the evolution of morality. His following book The Extended Phenotype: The Long Reach of the Gene (1999) is also very good, though more technical. Also excellent, if perhaps somewhat cynical about morality, is:
• Wright, R. (1994). The moral animal: The new science of evolutionary psychology. Pantheon Books.
The two standard textbooks on evolutionary psychology, both regularly updated:
• Workman, L. & Reader, W. (2026). Evolutionary psychology: An introduction (5th ed.). Cambridge University Press.
• Buss, D. M. (2025). Evolutionary psychology: The new science of the mind (7th ed.). Routledge.
A short list of refutations of common misconceptions about evolution, covering many of the motivations for attacks on evolutionary psychology (and on its predecessor sociobiology) by non-biologists:
• Al-Shawaf, L., Zreik, K., & Buss, D. M. (2021). Thirteen misunderstandings about natural selection. In T. K. Shackelford & V. A. Weekes-Shackelford (Eds.), Encyclopedia of evolutionary psychological science (pp. 8162–8174). Springer. https://labs.la.utexas.edu/buss/files/2018/07/13-Misunderstandings-About-Natural-Selection.pdf
A short theoretical online manifesto, from the Santa Barbara school (which emphasizes narrow task-specific modular adaptations more than most):
• Cosmides, L., and Tooby, J. (1997). Evolutionary psychology: A primer. Center for Evolutionary Psychology. https://www.cep.ucsb.edu/primer.html
A broad synthesis of evolutionary psychology with (in places a little dated) computational theory of mind:
• Pinker, S. (1997). How the Mind Works. W. W. Norton & Co..
A more philosophical viewpoint written by a professional philosopher very familiar with evolutionary psychology:
• Joyce, R. (2006). The evolution of morality. MIT Press. https://doi.org/10.7551/mitpress/2880.001.0001
Evolutionary explanations of ethics often feel a lot like “group selection” and succumb to the same general problem that individual selection (and kin selection) work faster and are more powerful than group selection
Actually, that problem has basically been solved in the last few decades. Briefly, morality can be explained via only individual/kin selection arguments. Collaborating with others in positive sum games (e.g iterated Prisoner’s Dilemma or Stag Hunt) is advantageous (most social primates can do this). Potential partners to cooperate with you have a good evolutionary reason to want to know that you’re a reliable partner before cooperating with you. Thus it’s to your advantage to have a good reputation as a reliable partner. From that, given humans evolved complex language, things like gossip, reputation, and justice then evolve. No group selection required.
More generally, evolutionary psychology has been researching this sort of question (with somewhat simplified mathematical frameworks — basically, looking at the question of what sorts of behaviors are likely to have higher eigenvalues in low rank approximations) for the last half century, and has made significant progress.
So, where did, for example, justice[1] come from?
Actually, we have a clue on this one. In experiments, humans generally (the majority of the time) show willingness to actively go out of their way to punish other people who have broken moral norms, even if neither they nor any of their family or friends were among the ones harmed by this. Psychologist call this “third-party punishment”, but this is fundamentally what the word “justice” means: enforcing moral norms for motivations other than direct personal advantage or revenge.
So far, in similar experiments on other primates, we haven’t been able to reproduce this behavior. So it appears justice evolved some time in the last ~ 8 million years since the last common ancestor between humans and chimps/bonobos, and is a trait unique to humans (along with language, gossip, reputation, and so forth).
The evolutionary theory is that humans’ willingness to participate in third-party punishment alters the incentives in the society to make breaking moral norms less rewarding, thus discouraging norm-breaking, so making society lower in crime, which is to the advantage of everyone doing this (and their friends and family). This is further encouraged by the fact that being known to participate in this way gives one a reputation for being moral and trustworthy oneself (“a fine upstanding citizen”), making others more willing to trust and interact on cooperative terms with you.
The humans still control the resources and determine the course of events, somehow, and use the universe mostly for their own purposes.
Or at least, the society made up of humans plus many AI minds much smarter than humans still has feedback systems that make it stably want and do what is actually best for the humans. I.e. the structure and enforcement mechanisms of the society ensures the ASI consensus is humanitarian. This requires the goals structure of the ASIs to have some rather different properties than many forms of goal-maximization would produce.
Try thinking of the situation like a primatologist. Preferably one with expertise in high status apes throwing shit at each other.
If should be trivially simple to recognize if model pretraining runs have been chained togther, looking just at the perplexity graph or hyperparameter schedule settings.
Running multiple pretraining runs in parallel is the same problem as distributed training, which has quite high bandwidth requirements between locations, so this should also be readily apparent with only quite basic monitoring of incoming/outgoing bandwidth.
However, doing this sort of stunt for RL post training runs is a lot more feasible.