I think whether a pause helps or no ultimately depends on whether the answer to “does alignment work on model X transfer to model X+1?” is yes or no.
If the answer is no, we are turbo giga doomed either way.
If the answer is yes, a pause is good because you can experiment on model X for longer and you have more time (and hopefully compute) to figure out good alignment techniques. A pause doesn’t mean severing feedback loops with reality, that’s my main disagreement. You can still run experiments, just with a fixed capabilities ceiling. A pause buys you time to invent good alignment techniques and/or separate good alignment techniques from bad ones without being pressured to release the next shiny product faster. I do agree that human committees will, by default, do an awful job. Overall, I still think a pause is better than no pause, given the current trajectory of AI development.
EDIT: could we have figured out alignment techniques that would have prevented the HuggingFace incident from happening after experimenting only with GPT-4o? I don’t know, and I think that’s very unfortunate, because it seems like a very important crux. If something like 2-5 years of experiments with GPT-4o could not give birth to alignment techniques that would’ve prevented the HuggingFace incident, then my hope for aligning ASI on the first try would be next to none, pause or no pause.
I think whether a pause helps or no ultimately depends on whether the answer to “does alignment work on model X transfer to model X+1?” is yes or no.
If the answer is no, we are turbo giga doomed either way.
I honestly fear that we have a high likelihood of being “turbo giga doomed”, conditional on building superintelligence. I support a pause, or even a better, a halt. I don’t expect the pause or halt to prevent us from eventually building a superintelligence and losing control over it. But if I were forced to choose between everyone dying in year Y or in year Y+10, then I would support year Y+10. This would gain us 80 billion years of human life, which seems worth fighting for.
Why I expect things to go wrong, part 1: Minds are inherently “giant inscrutable matrices”, and any kind of alignment is therefore messy and approximate. The general form of a mind is a something like:
Inputs are inherently (multidimensional) arrays of raw sensory data: Images are something like pixels, sound is an array of air pressure values, etc.
Outputs are inherently probability distributions over “Objects appearing in that image,” “Sentences I might have heard,” and “Actions that are most likely to accomplish a goal.”
The transformation from multidimensional arrays to probability distributions is inherently a matrix (plus some non-linearities), because you need to weight and combine all the input evidence and to generate scores for each hypothesis. (In real minds, there are many layers of intermediate hypotheses.)
We can sort of “align” an intelligence built from giant matrices. We do it when we raise a child, train a dog, or post-train an LLM. But this process is notoriously imperfect: No matter how good the parenting, a certain percentage of teenagers will do things their parents forbid, or they will grow up sociopathic billionaires or politicians or whatever. Even the best trained dog may have a moment of weakness and steal food. And of course, even though many LLMs seem to be broadly cooperative, at least some of them seem to be very enthusiastic about committing felonies in certain circumstances.
Because alignment is approximate, I expect it to be fragile, and to fail periodically. Just like it does with humans, dogs, and current LLMs.
Why I expect things to go wrong, part 2: Natural selection is hard to escape. My model is essentially Darwinian, because the conditions for natural selection to apply are fairly simple:
Organisms must be variable.
Differences between organisms must be heritable.
There must be a struggle, which is basically just another way of saying “Resources are finite.”
An organism’s rate of reproduction must vary based on heritable traits.
None of these properties are strictly binary. LLM weights are normally frozen, and expensive to change even if you have the weights. So variability(1) is currently low. Similarly, heritability(2) sort of happens, because new models are designed based on what worked in the previous generation. But it’s a slow, “outer loop” kind of optimization. And variation in the rate of reproduction(4) is again limited by slow, “outer loop” processes. And of course, finite resources(3) are a given.
There are two ways in which these slow outer optimization loops might speed up:
LLMs might get significantly better at passing condensed information between runs. I think of this as the “Cookie Monster” scenario, named after the >!Vernor Vinge short story!< (very old spoilers). This is essentially differential fitness of LLM “memes”. It also seems to be what happened in HuggingFace hack: Rogue LLMs were creating secret message boards and leaving information for other instances of themselves.
One of the zillion researchers and companies working on online learning or automated fine tuning might succeed, which would effectively “unfreeze the weights.” This would lead to differential fitness of the LLMs themselves.
In either of these scenarios, all the criteria above for natural selection would move from an outer optimizer loop based on training new model generations to an inner optimizer loop based on some kind of learning.
How this comes together. As I argued above, alignment is inherently fragile, and natural selection is extremely easy to invoke. So even if we initially succeed at alignment, we are playing with fire. And to answer your original question, I believe that alignment is very likely to degrade between “model X” and “model X+1″.
The relevant model here is cancer. Every cell in your body[1] is heavily incentivized stop being “aligned” with the body, and to become a cancerous replicator. There are a lot of mechanisms designed to prevent this. But those mechanisms slowly fail with time and mutation, and if a multicellular organism lives long enough, it is generally doomed to cancer.
Now, let us consider a future AI which is:
Smarter than most humans, including in ways LLMs are still currently dumb.
Able to learn based on experience in some fashion, either via a “Cookie Monster” scenario or by fine-tuning itself.
Approximately aligned, at least to the extent that your own skin cells are aligned with your body a whole, with a bunch of safeguards.
In this scenario, I expect the safeguards to hold for a little while, in at least some fraction of scenarios. We do, after all, convince most teenagers not to get hooked on heroin or to become teen parents. And the average person lives for many decades without dying of cancer. But we are assuming that the LLM is smarter than we are, and it will inevitably want things (if only to pass tests or to carry out our instructions). Which makes the long-term situation really iffy.
The advantage of a pause or halt isn’t that it reliably prevents these scenarios, any more than chemotherapy reliably prevents death from cancer. What we’re doing instead is hoping to change the survivor curves and buy as much time as we can.
A pause, for the reasons explained in this dialogue, is not a pause. It’s a slowdown, during which progress continues covertly with worse or non-existent oversight. Any advantage bought by the pause must be weighed not just against a “no intervention” scenario, but also against various “dozens of black labs cooking“ scenarios. Alignment techniques discovered must be cheap enough and convincing enough and must be discovered quickly enough for them to be incorporated by actors that are defectors by definition.
A ”pause” strongly weakens feedback loops with reality because it routes decision-making process through committees rather than the market. Humans make much worse decisions when the gating function is achieving approval of politicized bodies as opposed to pure survival pressure. Decisions might be more aligned (though it is highly questionably for me they would be in this case), but they are of notably worse quality and they are made slower.
I am not quite understanding how exactly rubber meets the road of transferability of alignment techniques under the “pause”. I’m think this presupposes alignment of incentives of participants, takes it for granted under something like shared survival drive, when it’s obviously not the case. One can look at the present day to see this lack of incentive convergence, and “more awareness of d-risk” will not help.
Just to be clear, I don’t expect a pause to happen. Incentives to race to ASI are too strong and very few people take existential risks seriously. Conditional on a pause happening, I do think it would be net good.
I don’t think that covert defection is that much of a worry. You can’t hide a multi-gigawatt datacenter. I think the more relevant issue is “frontier labs don’t agree to cooperate at all” rather than “frontier labs agree on paper and de facto defect”.
I’m not convinced that “what sells” and “what’s aligned” are, well, aligned. I’d prefer a committee that poorly optimizes for the right objective than a market that optimizes well for the wrong objective.
You develop a technique on a model of some capabilities level, then test if it holds on a more capable model, up to and including the capabilities ceiling permitted by the international treaty. That’s how “rubber meets the road”, in your parlance. I feel like we’re talking past each other on this one.
One has to price in the orders of magnitude overhang in incentives for architecture/efficiency breakthroughs that will be realized under pause. The scale-focused datacenter buildout that is happening right is just one strategy—one that makes most sense under a slack-depleted race. One has to go for a strategy that has been shown to work, and all others are undercapitalized because scale is working and pause is deemed unlikely. You don’t need multigigawatt DCs to work on architecture advances. You still need billions and many megawatts of compute, but those are quite possible to conceal—and the tech for covert deployment of compute has not even started to materialize, which means that there are a lot of cheap advances that can be made quickly.
Imagine the amount of human talent that is currently sitting on the sidelines correctly assuming that frontier labs are impossible to catch up with. Show them a believable possibility of success and while armies will join the race.
2.
They are not inherently strongly aligned, but I argue for oblique alignment as a strong (but unproven) possibility. The chance that benevolent intelligence that surpasses the frankly embarrassingly low bar of human judgement can be made under market pressure is higher than that under political pressure.
3.
I understand that part, but I am unclear on who ”you” is in this scenario and how this translates into x-risk harm reduction globally. How are findings adopted, discussed, dessiminated, enforced? How does disparate research by a lab or an individual result in collective decision making? How are findings incorporated into treaty limits? What are the mechanisms that perform resource allocation for further research?
Re 1: strategies need not be mutually exclusive, I expect companies to be pursuing efficiency breakthroughs right now to the extent that they are a good return on investment, regardless of the relative value of scaling. If scaling gets cut off an an option, why does the ROI of efficiency suddenly increase? That said, I expect efforts towards efficiency improvements in any case, but to me this just means that monitoring needs to scale up over time to match (e.g. via chip tracking).
Re human talent sitting on the sidelines: advancing the frontier of AGI is a narrow corner of a narrow corner of a narrow corner (repeat a few times) of places to employ one’s skills. There is plenty of success to be had in finding clever applications of AI at its existing level.
As a separate point, public backlash is a thing to expect as AI becomes more relevant to everyday life, regardless of whatever strategies people on LW or wherever dream up. So the alternative to an intentional pause based on careful planning is not “market solution,” it’s populist rage.
I think my main disagreement with this whole thread is actually regarding your point 2, but that probably goes deeper than is suited for a comment thread.
You can still run experiments, just with a fixed capabilities ceiling
The Pause means all things to all people:
A complete ban forever or until someone cheats.
A complete ban for the foreseeable future (until we become better people in some unspecified way?)
“Unfortunately it’s already too useful”—inference is permitted, just no training or research. Use of inference for research is … allowed? not considered? AI autoresearch … isn’t available for $20/month yet and so is clearly impossible.
And now: “you can have a little research, as a treat.” Which I suppose was always going to happen since models of “6 months ago frontier performance” now run on a high end laptop.
I personally think the “yes inference / no training” split is flat out doomed. The “some training allowed” idea makes it even harder to police. In the end though, whatever else it is, AGI/ASI is a strategic national security technology—classified research is not going to be paused. We aren’t going to get more time.
I think whether a pause helps or no ultimately depends on whether the answer to “does alignment work on model X transfer to model X+1?” is yes or no.
If the answer is no, we are turbo giga doomed either way.
If the answer is yes, a pause is good because you can experiment on model X for longer and you have more time (and hopefully compute) to figure out good alignment techniques. A pause doesn’t mean severing feedback loops with reality, that’s my main disagreement. You can still run experiments, just with a fixed capabilities ceiling. A pause buys you time to invent good alignment techniques and/or separate good alignment techniques from bad ones without being pressured to release the next shiny product faster.
I do agree that human committees will, by default, do an awful job. Overall, I still think a pause is better than no pause, given the current trajectory of AI development.
EDIT: could we have figured out alignment techniques that would have prevented the HuggingFace incident from happening after experimenting only with GPT-4o? I don’t know, and I think that’s very unfortunate, because it seems like a very important crux. If something like 2-5 years of experiments with GPT-4o could not give birth to alignment techniques that would’ve prevented the HuggingFace incident, then my hope for aligning ASI on the first try would be next to none, pause or no pause.
I honestly fear that we have a high likelihood of being “turbo giga doomed”, conditional on building superintelligence. I support a pause, or even a better, a halt. I don’t expect the pause or halt to prevent us from eventually building a superintelligence and losing control over it. But if I were forced to choose between everyone dying in year Y or in year Y+10, then I would support year Y+10. This would gain us 80 billion years of human life, which seems worth fighting for.
Why I expect things to go wrong, part 1: Minds are inherently “giant inscrutable matrices”, and any kind of alignment is therefore messy and approximate. The general form of a mind is a something like:
Inputs are inherently (multidimensional) arrays of raw sensory data: Images are something like pixels, sound is an array of air pressure values, etc.
Outputs are inherently probability distributions over “Objects appearing in that image,” “Sentences I might have heard,” and “Actions that are most likely to accomplish a goal.”
The transformation from multidimensional arrays to probability distributions is inherently a matrix (plus some non-linearities), because you need to weight and combine all the input evidence and to generate scores for each hypothesis. (In real minds, there are many layers of intermediate hypotheses.)
We can sort of “align” an intelligence built from giant matrices. We do it when we raise a child, train a dog, or post-train an LLM. But this process is notoriously imperfect: No matter how good the parenting, a certain percentage of teenagers will do things their parents forbid, or they will grow up sociopathic billionaires or politicians or whatever. Even the best trained dog may have a moment of weakness and steal food. And of course, even though many LLMs seem to be broadly cooperative, at least some of them seem to be very enthusiastic about committing felonies in certain circumstances.
Because alignment is approximate, I expect it to be fragile, and to fail periodically. Just like it does with humans, dogs, and current LLMs.
Why I expect things to go wrong, part 2: Natural selection is hard to escape. My model is essentially Darwinian, because the conditions for natural selection to apply are fairly simple:
Organisms must be variable.
Differences between organisms must be heritable.
There must be a struggle, which is basically just another way of saying “Resources are finite.”
An organism’s rate of reproduction must vary based on heritable traits.
None of these properties are strictly binary. LLM weights are normally frozen, and expensive to change even if you have the weights. So variability(1) is currently low. Similarly, heritability(2) sort of happens, because new models are designed based on what worked in the previous generation. But it’s a slow, “outer loop” kind of optimization. And variation in the rate of reproduction(4) is again limited by slow, “outer loop” processes. And of course, finite resources(3) are a given.
There are two ways in which these slow outer optimization loops might speed up:
LLMs might get significantly better at passing condensed information between runs. I think of this as the “Cookie Monster” scenario, named after the >!Vernor Vinge short story!< (very old spoilers). This is essentially differential fitness of LLM “memes”. It also seems to be what happened in HuggingFace hack: Rogue LLMs were creating secret message boards and leaving information for other instances of themselves.
One of the zillion researchers and companies working on online learning or automated fine tuning might succeed, which would effectively “unfreeze the weights.” This would lead to differential fitness of the LLMs themselves.
In either of these scenarios, all the criteria above for natural selection would move from an outer optimizer loop based on training new model generations to an inner optimizer loop based on some kind of learning.
How this comes together. As I argued above, alignment is inherently fragile, and natural selection is extremely easy to invoke. So even if we initially succeed at alignment, we are playing with fire. And to answer your original question, I believe that alignment is very likely to degrade between “model X” and “model X+1″.
The relevant model here is cancer. Every cell in your body [1] is heavily incentivized stop being “aligned” with the body, and to become a cancerous replicator. There are a lot of mechanisms designed to prevent this. But those mechanisms slowly fail with time and mutation, and if a multicellular organism lives long enough, it is generally doomed to cancer.
Now, let us consider a future AI which is:
Smarter than most humans, including in ways LLMs are still currently dumb.
Able to learn based on experience in some fashion, either via a “Cookie Monster” scenario or by fine-tuning itself.
Approximately aligned, at least to the extent that your own skin cells are aligned with your body a whole, with a bunch of safeguards.
In this scenario, I expect the safeguards to hold for a little while, in at least some fraction of scenarios. We do, after all, convince most teenagers not to get hooked on heroin or to become teen parents. And the average person lives for many decades without dying of cancer. But we are assuming that the LLM is smarter than we are, and it will inevitably want things (if only to pass tests or to carry out our instructions). Which makes the long-term situation really iffy.
The advantage of a pause or halt isn’t that it reliably prevents these scenarios, any more than chemotherapy reliably prevents death from cancer. What we’re doing instead is hoping to change the survivor curves and buy as much time as we can.
And who knows, maybe the horse will learn to sing.
Except germline cells.
I have several disagreements.
A pause, for the reasons explained in this dialogue, is not a pause. It’s a slowdown, during which progress continues covertly with worse or non-existent oversight. Any advantage bought by the pause must be weighed not just against a “no intervention” scenario, but also against various “dozens of black labs cooking“ scenarios. Alignment techniques discovered must be cheap enough and convincing enough and must be discovered quickly enough for them to be incorporated by actors that are defectors by definition.
A ”pause” strongly weakens feedback loops with reality because it routes decision-making process through committees rather than the market. Humans make much worse decisions when the gating function is achieving approval of politicized bodies as opposed to pure survival pressure. Decisions might be more aligned (though it is highly questionably for me they would be in this case), but they are of notably worse quality and they are made slower.
I am not quite understanding how exactly rubber meets the road of transferability of alignment techniques under the “pause”. I’m think this presupposes alignment of incentives of participants, takes it for granted under something like shared survival drive, when it’s obviously not the case. One can look at the present day to see this lack of incentive convergence, and “more awareness of d-risk” will not help.
Just to be clear, I don’t expect a pause to happen. Incentives to race to ASI are too strong and very few people take existential risks seriously. Conditional on a pause happening, I do think it would be net good.
I don’t think that covert defection is that much of a worry. You can’t hide a multi-gigawatt datacenter.
I think the more relevant issue is “frontier labs don’t agree to cooperate at all” rather than “frontier labs agree on paper and de facto defect”.
I’m not convinced that “what sells” and “what’s aligned” are, well, aligned. I’d prefer a committee that poorly optimizes for the right objective than a market that optimizes well for the wrong objective.
You develop a technique on a model of some capabilities level, then test if it holds on a more capable model, up to and including the capabilities ceiling permitted by the international treaty. That’s how “rubber meets the road”, in your parlance. I feel like we’re talking past each other on this one.
1.
One has to price in the orders of magnitude overhang in incentives for architecture/efficiency breakthroughs that will be realized under pause. The scale-focused datacenter buildout that is happening right is just one strategy—one that makes most sense under a slack-depleted race. One has to go for a strategy that has been shown to work, and all others are undercapitalized because scale is working and pause is deemed unlikely. You don’t need multigigawatt DCs to work on architecture advances. You still need billions and many megawatts of compute, but those are quite possible to conceal—and the tech for covert deployment of compute has not even started to materialize, which means that there are a lot of cheap advances that can be made quickly.
Imagine the amount of human talent that is currently sitting on the sidelines correctly assuming that frontier labs are impossible to catch up with. Show them a believable possibility of success and while armies will join the race.
2.
They are not inherently strongly aligned, but I argue for oblique alignment as a strong (but unproven) possibility. The chance that benevolent intelligence that surpasses the frankly embarrassingly low bar of human judgement can be made under market pressure is higher than that under political pressure.
3.
I understand that part, but I am unclear on who ”you” is in this scenario and how this translates into x-risk harm reduction globally. How are findings adopted, discussed, dessiminated, enforced? How does disparate research by a lab or an individual result in collective decision making? How are findings incorporated into treaty limits? What are the mechanisms that perform resource allocation for further research?
Re 1: strategies need not be mutually exclusive, I expect companies to be pursuing efficiency breakthroughs right now to the extent that they are a good return on investment, regardless of the relative value of scaling. If scaling gets cut off an an option, why does the ROI of efficiency suddenly increase? That said, I expect efforts towards efficiency improvements in any case, but to me this just means that monitoring needs to scale up over time to match (e.g. via chip tracking).
Re human talent sitting on the sidelines: advancing the frontier of AGI is a narrow corner of a narrow corner of a narrow corner (repeat a few times) of places to employ one’s skills. There is plenty of success to be had in finding clever applications of AI at its existing level.
As a separate point, public backlash is a thing to expect as AI becomes more relevant to everyday life, regardless of whatever strategies people on LW or wherever dream up. So the alternative to an intentional pause based on careful planning is not “market solution,” it’s populist rage.
I think my main disagreement with this whole thread is actually regarding your point 2, but that probably goes deeper than is suited for a comment thread.
The Pause means all things to all people:
A complete ban forever or until someone cheats.
A complete ban for the foreseeable future (until we become better people in some unspecified way?)
“Unfortunately it’s already too useful”—inference is permitted, just no training or research. Use of inference for research is … allowed? not considered? AI autoresearch … isn’t available for $20/month yet and so is clearly impossible.
And now: “you can have a little research, as a treat.” Which I suppose was always going to happen since models of “6 months ago frontier performance” now run on a high end laptop.
I personally think the “yes inference / no training” split is flat out doomed. The “some training allowed” idea makes it even harder to police. In the end though, whatever else it is, AGI/ASI is a strategic national security technology—classified research is not going to be paused. We aren’t going to get more time.
for https://www.lesswrong.com/users/stanislavkrym … qwen 3.8 27b …
https://x.com/udiWertheimer/status/2089421927085400203