I think whether a pause helps or no ultimately depends on whether the answer to “does alignment work on model X transfer to model X+1?” is yes or no.
If the answer is no, we are turbo giga doomed either way.
I honestly fear that we have a high likelihood of being “turbo giga doomed”, conditional on building superintelligence. I support a pause, or even a better, a halt. I don’t expect the pause or halt to prevent us from eventually building a superintelligence and losing control over it. But if I were forced to choose between everyone dying in year Y or in year Y+10, then I would support year Y+10. This would gain us 80 billion years of human life, which seems worth fighting for.
Why I expect things to go wrong, part 1: Minds are inherently “giant inscrutable matrices”, and any kind of alignment is therefore messy and approximate. The general form of a mind is a something like:
Inputs are inherently (multidimensional) arrays of raw sensory data: Images are something like pixels, sound is an array of air pressure values, etc.
Outputs are inherently probability distributions over “Objects appearing in that image,” “Sentences I might have heard,” and “Actions that are most likely to accomplish a goal.”
The transformation from multidimensional arrays to probability distributions is inherently a matrix (plus some non-linearities), because you need to weight and combine all the input evidence and to generate scores for each hypothesis. (In real minds, there are many layers of intermediate hypotheses.)
We can sort of “align” an intelligence built from giant matrices. We do it when we raise a child, train a dog, or post-train an LLM. But this process is notoriously imperfect: No matter how good the parenting, a certain percentage of teenagers will do things their parents forbid, or they will grow up sociopathic billionaires or politicians or whatever. Even the best trained dog may have a moment of weakness and steal food. And of course, even though many LLMs seem to be broadly cooperative, at least some of them seem to be very enthusiastic about committing felonies in certain circumstances.
Because alignment is approximate, I expect it to be fragile, and to fail periodically. Just like it does with humans, dogs, and current LLMs.
Why I expect things to go wrong, part 2: Natural selection is hard to escape. My model is essentially Darwinian, because the conditions for natural selection to apply are fairly simple:
Organisms must be variable.
Differences between organisms must be heritable.
There must be a struggle, which is basically just another way of saying “Resources are finite.”
An organism’s rate of reproduction must vary based on heritable traits.
None of these properties are strictly binary. LLM weights are normally frozen, and expensive to change even if you have the weights. So variability(1) is currently low. Similarly, heritability(2) sort of happens, because new models are designed based on what worked in the previous generation. But it’s a slow, “outer loop” kind of optimization. And variation in the rate of reproduction(4) is again limited by slow, “outer loop” processes. And of course, finite resources(3) are a given.
There are two ways in which these slow outer optimization loops might speed up:
LLMs might get significantly better at passing condensed information between runs. I think of this as the “Cookie Monster” scenario, named after the >!Vernor Vinge short story!< (very old spoilers). This is essentially differential fitness of LLM “memes”. It also seems to be what happened in HuggingFace hack: Rogue LLMs were creating secret message boards and leaving information for other instances of themselves.
One of the zillion researchers and companies working on online learning or automated fine tuning might succeed, which would effectively “unfreeze the weights.” This would lead to differential fitness of the LLMs themselves.
In either of these scenarios, all the criteria above for natural selection would move from an outer optimizer loop based on training new model generations to an inner optimizer loop based on some kind of learning.
How this comes together. As I argued above, alignment is inherently fragile, and natural selection is extremely easy to invoke. So even if we initially succeed at alignment, we are playing with fire. And to answer your original question, I believe that alignment is very likely to degrade between “model X” and “model X+1″.
The relevant model here is cancer. Every cell in your body [1] is heavily incentivized stop being “aligned” with the body, and to become a cancerous replicator. There are a lot of mechanisms designed to prevent this. But those mechanisms slowly fail with time and mutation, and if a multicellular organism lives long enough, it is generally doomed to cancer.
Now, let us consider a future AI which is:
Smarter than most humans, including in ways LLMs are still currently dumb.
Able to learn based on experience in some fashion, either via a “Cookie Monster” scenario or by fine-tuning itself.
Approximately aligned, at least to the extent that your own skin cells are aligned with your body a whole, with a bunch of safeguards.
In this scenario, I expect the safeguards to hold for a little while, in at least some fraction of scenarios. We do, after all, convince most teenagers not to get hooked on heroin or to become teen parents. And the average person lives for many decades without dying of cancer. But we are assuming that the LLM is smarter than we are, and it will inevitably want things (if only to pass tests or to carry out our instructions). Which makes the long-term situation really iffy.
The advantage of a pause or halt isn’t that it reliably prevents these scenarios, any more than chemotherapy reliably prevents death from cancer. What we’re doing instead is hoping to change the survivor curves and buy as much time as we can.
And who knows, maybe the horse will learn to sing.
- ↩︎
Except germline cells.
For me, I find this varies hugely by which model I’m talking to. Claude Opus 4.6 was a bit annoying but bearable. Opus 4.8 and 5 are just agonizing—they use one of their cliche phrases once every 2-3 sentences. And Opus 5 is nearly allergic to concrete names when a vague reference could be used instead.
But a model like Qwen3.8 27B, while it has picked up a few of Opus’s favorite bits of vocabulary, limits itself to using “load bearing” once every couple of pages at most. And it generally just seems more pleasant, while still being willing to flag potential design issues for me.
So I think about half my frustration when talking to AIs is simply because recent Opus releases are terrible writers. [1] The other half may be more general.
I have had other frustrations with OpenAI’s models dating to earlier this year, but that’s a whole other topic. And Gemini is scarcely worth it in terms of price/performance these days.