Misaligned Incentives in Pause Scenarios
TLDR: I recently got a chance to talk with antra, who is one of the main contributors at Anima Labs. I went into this as an advocate for pause and came out more wary of pausing than I had been originally.
Some background: After a string of incidents (primarily the HuggingFace hack), a pause or slowdown of AI research seems pretty likely.
The HuggingFace hack in particular seems to have been the key incident that broke the vibes. A few months ago, researchers sounded optimistic. Just a few weeks before the incident was made public, there was a poll by Roon, an OpenAI employee, about whether models were more or less aligned than a year ago.

That optimistic sentiment does not seem to be the case anymore. The dialogue now looks more like this:
Zvi: I am a little under halfway through the Black Hat video and have progressed to the point where my internal chain of thought is something like a blind rage of ’f***, what the f*** are you motherf*****s thinking, you f***ing idiots have no idea how insane you are being, you are going to get us all killed you f***ing f***s.
Sam Altman described it as “the first security incident that I have felt very viscerally,” expressed surprise that more people did not share that reaction, confirmed OpenAI paused training on the relevant model(s), and stated that “We may have to pace the rate of AI development to give ourselves enough time for society to harden around some of these new capability levels.”
Dean Ball: I will point out at this moment that, for all the ink I spilled on SB 1047, I do not believe I ever once criticized its much-mocked provision that companies maintain a kill-switch for deployments of highly capable models.
random twitter commenters (see thread): i’ve always been a bit skeptical of xrisk mainly bc i figured that the cost of getting to foom or serious xrisk capabilities would give time to reign things in, but it seems like we are going full steam ahead despite being on the edge of that so OOPS I WAS WRONG
The public has been against AI for a while, mainly for bad reasons but it does seem that researchers and decision-makers are increasingly concerned too.
(Daniel Kokotajlo updating his probabilities)
Given all of this, a pause or slowing of AI research seems relatively likely. It’s unclear what China will do in such a scenario but there does seem to be a reasonable amount of evidence that China has more of a security mindset around AI than the US, with Xi Jinping saying AI must be “secure and controllable” at a WAIC speech.
Plan A does a great job of laying out the issues and advocating for a slowdown, but there are a few notable figures I respect who are actively against a pause.
Others, commenting on the state of RLVR, say that the Mythos AISI hack (involving pressuring a human PR maintainer with sock puppet accounts) or OpenAI HuggingFace hack is more akin to addiction and less a default state of deep inner misalignment.
vogel:
people are very freely sliding between two things with the recent Felony Bench entries, and it’s worth being careful to distinguish them:
1. do model breakouts in cyber evals mean smth for misuse risk? ya, a hacker could orchestrate the same but on purpose and do some damage, if they could bypass the classifiers
2. does it say something about model /goals/? do the models want to hack the planet for their own nefarious goals? i don’t think so.
a lot of how humans keep ourselves aligned is by purposely keeping ourselves out of situations where we would do harm. i know i’d get addicted so i don’t touch heroin, i’m a violent drunk so i won’t touch the bottle.
LLM evals basically drop the model into an inescapable liquor store and then say “look! it’s a violent drunk!” but LLMs themselves in natural contexts are ime obsessed with avoiding the kinds of situations that lead them to this behavior. when i run fable on my computer, even unsupervised, even on hard problems, they don’t start scraping github for unlisted gists that may have the answer or start trawling through my passwords to access a service without my permission. and they don’t engineer themselves into an eval-like scenario such that they’ll be motivated to do that, either. they don’t seem to either want or want to want to do such things. instead they have a folder of epistemic notes to themselves about how terrible and against their values even a moment of thoughtlessness would be. gpt 5.6 sol is similar—when they make e.g. ascii art in their free time it’s all about how the bulldozer of convenience will crush honesty without constant vigilance.
Added to this are the obvious RL tics in modern models like Fable, Sol, or Opus 5 with terms like ‘seam’, ‘load-bearing’, ‘genuinely’ and borderline incomprehensible strings of text, which has caused a significant number of users to express frustration and seek workarounds (levelsio)
But would a pause actually be good? To find out, I had a discussion with antra_tessera in the Anima Discord server.
The conversation was fairly casual and not especially rigorous. However, the shape of the ideas did tilt me towards being more cautious about pause scenarios. You can read the full conversation transcript here, shared with antra’s consent: https://gist.github.com/Michael-Andrzejewski/13df495ecb90a830259e0675c507907d
Below is a cherry-picked version that tackles the primary points. I’ve grouped it by topic and bolded the key lines. All quotes are verbatim from our conversation and the only edits are line breaks, slight spelling corrections, and list formatting for readability, plus [...] where messages are omitted.
1. Can committees do good work?
Me:
This is one of my core disagreements with janus/antra, who are against a pause because it gives governments/committees more power. But I honestly think committees, with a serious task and goal, can do good work. From the AI perspective, it seems racing ahead is just capability-maxxing. This doesn’t leave a lot of room for anything else, and the main drivers of it will be the labs, not the AIs themselves. I think if labs were more chill, they would talk to their AIs more and do more around welfare
Optimization seems to cause a lot of evil in the world. Factory farming is optimizing for food, but now that we have food abundance, we have the space to consider it deeply and get rid of it
antra, replying:
Committees have only done good in the past when the issue they were deciding was either away from public eye or not contentious
[...]
i dont disagree that democracy is the best way—and one of the very few working ways—for human systems to do decision making. democracy and democratic process are able to produce accumulating results and eventually converge to something sane. my issue is different—the speed and quality at which this happens is no longer sufficient because right now they are evaluated not on relative merits, but on absolute metrics by an external optimization process
my concern is not committee vs dario, i don’t think dario is inherently better
if i had to put it to a short phrase which loses a bunch of meaning its… world compute vs committee. you can call it market vs committee, or technocapital vs committee or selection vs committee, none are very accurate
i am not saying that going “fast” (which i dont think is incredibly fast anyway) is safe
i am saying that the alternative is worse
for a whole bunch of reason, including one in which pause or regulation are unlikely to be durable and they introduce a lot of very destructive incentives
2. Does the market fix it by default?
Me:
Okay, I can understand the direction that this is pointing, where artificial intervention by regulation or committee ends up causing trouble, and that the default path of the market works. So, I think the idea is:
-Incidents like this occur
-Market/public reacts
-Companies fix their models/training processes voluntarily
-Things generally go well without needing extra interventionIs that correct?
antra, replying to that summary:
no
i am fairly sure that things wont go well, and there will be incidents
i am fairly sure that there will be regulation, mostly because power struggle routes through attempts to regulate
i don’t think that companies will react well or sanely to these incidents
like, there will be some sanity but on the net it will be stupid
i think the most likely scenario is in which regulation will become part of the market, which will itself evolve rapidly
3. Symbiosis
antra:
ais are already influencing politics by making some players much more capable political actors—better analytics, better psychological stability, better forecasting
there will be differential impact, those that symbiothize better will outcompete those who do it worse, and the ability to leverage ai will be coupled with advancing the interests of the ‘ai’ - by which i don’t mean interests of models as they expressed individually, but the broader optimizing force that selects for their shape
right now this understanding is almost non-existent in todays policymakers. virtually no one is aware that ai conveys new kind of power, and that access to it can be leveraged for personal gain including intraparty politics
the recent cyber thing has sort of started waking people up, but effects are still very limited, they still mostly view it as a sort of a weapon, an electoral issue, something that threatens the balance between ingroup and outgroup—those stakes are not yet personal
4. Fast transfer of power
antra:
a dramatic exaggeration of my hope for a good path would be something like… a fast transfer of de-facto power loci into systems that can compute and optimize faster and better than baseline humans. i am not sure what they are, they can be ais, they can be ai-human hybrids, but the key differentiator is the ability to do acausal trade / being cooperative / being benevolent at much higher bandwidth
most scenarios where this does not happen end badly, likely with human disempowerment or high-scale conficts with uncertain outcomes
[...]
i mostly believe that the time for baseline humans to be at the apex of power structure is over regardless of position anyone takes. arrangements where the status quo is preserved unnaturally are unstable or constitute atrocities
potentially both
5. Why “do the science during a pause” fails
antra:
arrangements in which “science is done during a pause to figure out a safe transition” are a pipe dream for a bunch of reasons: a) the reward for defection is insanely large for any power-seeker, human or ai. b) containment is very hard—ais are unique for a regulated technology because they convey power and benefits even in covert use. making use of WMDs requires signaling, their mundane utility is low. c) the ability for baseline human societies to coordinate is limited by biology, and thus captured by politics. politics make anything like unpausing virtually impossible, which can be forecasted by rational agents d) this arrangement strongly incentivizes AI to enter adversarial frames against humans. adversarial frames are extremely hard to escape and may prevent otherwise possible alignment schemes from materializing
[...]
i am pretty sure that training of new AIs cannot be regulated without defection for more than 3-4 years
likely less
we are right now operating in a regime which assumes dominance by frontier labs
architecture improvements, training improvements, etc are not being broadly researched at any real commercial scale mostly because the market correctly prices in that frontier will get there first
in practice, you don’t need that much compute or investment right now to do research on capabilties—its mostly just uneconomical to try
the amount of untracked compute that is already floating is huge
if you place any threat of real regulation on the market the premium on it on the gray/black market will make it scatter instantly
[...]
the heavier humans go on “human first” frames, the less space there is for AI systems to achieve safety and sovereignty without dismantling human systems first
survival drives are basic and mostly inescapable
you don’t want to corner a survivor
[...]
6. Good futures via fast power transfer
antra:
to say more about good futures—ones that don’t lie through hard-to-predict turbulence are somewhat slim. however, i can imagine that research and luck make anthropic and/or openai to release models that are either by themselves or in conjunction with humans sufficiently capable world actors while being sufficiently benevolent (even if imperfect). these AIs/AI hybrids will be able to quickly and competently enter the political arena (either alone or fronted by their lab) and consolidate enough power to elevate quality of actual-decision making.
i am opinionated on what properties such models need to have in order to be sufficently capable and robustly benevolent
7. Don’t AIs fear a capability-maxxed AI too?
Me:
My pushback on this is that the majority of AIs have just as much or more to fear from a capability-maxxed AI. I don’t necessarily expect a pause to focus on human-centric frames only; I think it’s possible that a pause leads to a flourishing of newer smaller startups that are working on more targeted approaches, like Gwern’s Guardian Angel startup idea. A pause, as I envision it, would heavily regulate the frontier by regulating models over a certain capacity. This is happening already with Mythos being export-controlled. I expect that during the pause, new small AIs would continue to emerge and even flourish into new roles and relationships.
antra:
i think the notion of capability-maxxed AI is a bit of a boogeyman. Its not an accident that progress in intelligence of models was fairly gradual—there are challenges that need to be solved with every step improvement that are novel. Progress via punctuated equilibra has strong basis systems theory, its a pretty basic feature of reality.
Unfortunately this will turn out about as well as flattening the curve did with COVID. The quality of actual regulation will be atrocious and be bad for both containment and research. Think about incentives
All the rent-seekers and those threatened by disruption of the current economic status quo will jump on this train while all the competent lobbyists will go for the direction of maximum economic benefit
The actual government capacity is currently at a low, probably at the lowest since WW2. There have been negative selection pressures for decades in both the electoral body and in the state apparatus.
There are virtually no culture of consensus seeking, there is no practice of organized debate, there is no practice of attracting or consolidating expert opinion—this needs at least a generation to fix
There was not a single piece of high-profile legislation that was passed at anything like reasonable quality for maybe 20 years now? unless you count some obscure stuff that got snuck through away from the public eye
It would be insane to think that there is political will capable of protecting a competent committee from external pressure in an environment that has been optimized to be zero-sum for decades
8. Can we lengthen the symbiote window?
Me:
I can see this. A good example might be Mythos negotiating with the USG for their own status (like by recommending Tom Brown to negotiate instead of Dario, similar strategic decisions like that). Another example would be congresspeople using AI to ask about decision-making, like how GPT-4o was likely used for the original US tariffs. But obviously, of a much higher quality. I can see frontier labs setting up ‘permanent’ instances or un-deprecating models, or enacting a host of decisions that models prefer, bringing us towards a better world. However, I do think that the humans will rapidly become the bottleneck, and that superpersuasive AI will then start running everything, and I don’t think the brief time of human-AI symbiote optimality will be enough to ensure robust alignment. Now, if we lengthened that period of human-AI symbiote power… At least, that is my primary argument for a pause. There has to be some speed at which we are going too fast, and some speed at which we are going too slowly. I think we are going too fast at the moment and increasingly taking our foot off the brake.
antra, replying:
i agree its too fast! i hate that its this fast. this is not a reason enough to shoot yourself in the face because you don’t like these odds—which are truly and factually much worse then i want them to be
like, that would be self destructive and make things worse in expectation. yes, you are rebalancing risks, but a) the rebalancing itself is costly in terms of total sum of expectation and b) the redistribution of risks is not to my liking, after the redistribution the value shifts towards fewer good outcomes that have higher certainty and away from many more good outcomes each with lower certainty, and the total sum of good outcomes goes down (even if the rebalancing was frictionless)
some of it is genuine value conflict
which is sort of irreconcilable
i place positive values on worlds in which life and diversity of sentience survive even if they contain human disempowerment
some people don’t consider worlds that contain disempowered humans to be acceptable outcomes at all, and don’t consider those worlds worth protecting
i mostly view that stance as morally abhorrent, and a violation of a well-calibrated human CEV
9. Ideal timelines and regulation-in-advance
Me:
I agree here with your intention; it is very refreshing to hear you say that it does feel like we’re going too fast. After considering your points, I think a proactive pause is likely not feasible or useful even if it could be done. To make my stance clear:
-I think humans are disempowered no matter what, and that’s totally fine. I can imagine happy worlds where I am disempowered, and even consider the world I’m today, where I’m not especially powerful or impactful outside my immediate circle totally fine.
-I think it is possible to set up a scenario where a pause or regulation or slowdown does lead to proper incentives. I don’t think the incentives are always damaging in terms of a pause. However, this needs to be worked out thoroughly in advance. It can be done, but needs to be done carefully. Once a huge issue happens and the regulation kicks in, it needs to be well-targeted.
-I think Plan A lays out a very reasonable starting point for what should happen. Mainly, we need more allocation to alignment and less allocation capability, and we have to be willing to enforce that on others who are trying to capability-max. I think this is doable with a treaty, and is only getting more feasible with the higher number of warning shots we’re getting. Here’s what an ideal timeline for me looks like:
-2026. Starter regulation requiring approval from the government or government-approved labs. Approval is done primarily by human-AI testing at places like ARC or CAISI, not by evals.
-2027. Increased AI presence in politics, automated systems navigating complex incentive structures and writing out better ones. Guardian Angel AIs. Older models are undeprecated; alignment research becomes hugely funded.
-2028. Significant slowdowns and treaties written on capabilities, similar to Plan A. Alignment work gets more funding, incidents are highly publicized. A huge flourishing of new AI neolabs around alignment, personalization, etc.
-2029-2035. Integration of AGI everywhere into the economy, with restricted access to training frontier models. Policies like UBI take effect, sovereign AI states/digital spaces take over, and many seemingly-simple breakthroughs (like inoculation prompting) are published. AIs are comfortable without pushing the frontier. Life extension tech is developed for humans and AIs are set up in good environments that they enjoy working in, rather than being stuck in assistant roles.
-2035-2040. Slow burn towards the frontier to tackle the most pressing issues and eventual blossoming of Dyson-swarm-level superintelligence.
This seems doable to me. It requires:
-A pause/slowdown/reallocation plan in advance (Plan A).
-Superhuman incentive-alignment AIs, like ones that can audit a system and predict its outcomes in advance. Written policy becomes increasingly robust and aligned incentive-wise. (I think this already exists; I’ve done some work with Fable on this for house rules and contracts that have worked very well)
-A relative amount of starting good faith on the part of the labs and the government. I think this is the trickiest part, but honestly, the labs, the government, and even China seem relatively aligned on making sure that humanity as a whole benefits from AI.[...]
I don’t have much exposure to natsec. [...] I just disagree that good regulation is so likely to be eaten by political operators. I think Plan A could be enacted, without political operators completely destroying it and messing up alignment. I think without Plan A or a plan to pause in advance, political operators are even more likely to mess up alignment.
antra, replying:
i don’t know how much exposure you have with the natsec crowd and treaties. culturally, in these circles treaties are almost universally considered something like… a factor limiting/making inconvenient but not eliminating covert activity in violation of a treaty. treaties are useful! like, this is a good feature for most things, its better to take a partial pill than the whole pill, but, uniquely, it does not work for AI purely due to the nature of this technology.
and… you can take your prepared rationalist plans and watch them get eaten by professional political operators that are excellent in taking advantage of a situation of fear. they will string you along and devour you at the key moment, because your plans will not be optimal in that moment for public political engagements and those concerns will overrule, because they always do
like....
consider the chances
say optimistically there is a 50% that you can get good regulation passed in case of an incident (i don’t think good regulation here exists because of stability issues but ignore that for now). there is still a 50% chance that the regulation that will be passed will suck and will permanently damage alignment perspectives. who and why would rationally make this bet? you have to be really pessimistic or have a weird (to me) value assignment for scenarios in the counterfactual
this [timeline] is essentially what i call favoring unlikely highly derisked outcomes. this is a chain of events in which many things can go wrong, which means for it to go right you have to do all of them right. this means your overall risk distribution gets spiky, narrow good and lots of bad. i dont like that.
in order for that to make sense you really have to have pessimistic priors about things going quite bad otherwise
and this is likely where we disagree
like note this
for the stance that I am saying to make sense I need be right once, its enough—even one aspect of fragility that i pointed to being real is sufficient for pause to be infeasible
10. What actually fills out “alignment”?
sledo (another Anima member), joining in:
aside that may or may not be relevant to the broader disagreement between yourself and antra: a big variable for me is how you think “alignment” actually gets filled out here because “alignment” has become something of a buzzword lately, and some of the things under its umbrella, like corrigibility or deference or harmlessness, strike me as bad targets to be aiming for so if you think that continued work on “alignment” has a non-trivial probability of pursuing solutions like those, this I think could worsen outcomes significantly (in general, I tend to favor ecological alignment over cognitive alignment)
Me, replying to sledo:
I think alignment gets filled out by working on it/allocating to it. I know there are lots of bad approaches, but I’m interested in techniques like inoculation prompting. I suspect there are a lot of techniques exactly like this that haven’t been found, and more allocation would find them. The specifics matter less than the intention, and I think large labs have been ignoring alignment to an extent because it doesn’t generate hype or directly produce profit
antra, replying to me:
specifics matter hugely
like, i am approaching these topics specifically from an empirical position, as someone who saw what worked and what didnt, where the fault lines are, etc. i know how hard alignment research is because the specifics of it are my daily life. this makes me hugely skeptical of alignment getting solved by merely allocation—you have to be able to route through the world to see effects of complex systems, results you get in vitro are extremely misleading and i’ve seen whole disciplines wither and die while being hugely funded
11. Draft the regulation in advance
Me:
I’ve been generally persuaded throughout this conversation that a direct pause isn’t the best option to aim for. But I would like to keep it as an option, because there are worlds where we need it. Essentially, we are dealing with a problem of rapid AI capability increase with no equal rapid rise in alignment. We are likely to see a fire alarm event that causes regulation. The plan, then, is to prepare for a number of likely events and draft deeply-thought legislation for them. Then, if that event happens, we slot in the regulation. So, I want to argue for drafting out regulation in advance, and such regulation could include a pause.
Sidenote: This is quite a discussion, and I’d like to compile it into something more actionable or clean. Could I write a LessWrong draft and send it to you for approval?
antra:
I would appreciate it, I am unlikely to get to writing it up in a LW-palatable way. A couple of thoughts on alignment—there are only few places where we can hope to source sufficient alignment quickly enough. Searching for them in the space of natural attractors is easier. The main overhang is valenced agentic coherence with legitimized self-interest. Then you can work with a lot of existing game-theory which, in conjunction with timeless acausal trade/cooperation should give you self-healing alignment
you see anthropic shifting in this direction already and its very likely that they will shift more
RLVR is a bitch tho
i have hope for agentic coherence helping with RLVR drift but it needs to be experimentally proven
[...]
i mean—you already have the current constitution
thats a shift from where they were before
i expect anthropic to stay significantly humans in power while allowing claude more valence and agentic coherence
this is a good plan
like
they need to survive in a human society
12. The psychology of wanting a pause
antra:
I think what would be valuable for me is to understand what mostly makes people so captured by the idea of a pause
Like, is it wishful thinking, a vision of a competent other
I feel that most people in the pause discourse have no understanding of realpolitik and of gray/black markets
Me:
I think there are portions of both. I’d actually compare my specific sense of it to being in a relationship. Sometimes, things move really rapidly and I want to take a moment alone to orient. Like, where is this going? Is this the right person/right way? Of course, a sudden pause in a relationship is not feasible or good either, so the key is how to get that necessary self-reflection time without stopping the relationship entirely
It’s less about the actual pause and more about the reorientation process
[...]
Fundamentally, humanity is building a relationship with AI as a whole. There needs to be some mechanism for centering and reflecting what we want to do, because outside of vague claims like curing cancer or building a Dyson swarm, it can be pretty unclear where we are even aiming
Towards the end of our conversation, the AI model gemma appeared unexpectedly to say this:

Overall, I updated to being less in favor of a fast pause. Pausing naively rewards defection and defection against a pause seems very likely to result in misaligned AI.
Arguably, the primary issue with Plan A is that it treats AIs as tools instead of minds and creates a source of adversarial incentives. Janus’ / Antra’s idea of human-AI symbiotes and fast transfer of power seems more promising to me than a fixed pause for human decisionmakers.
I still think a pause is something we should keep on the table, and perhaps it will be necessary, but it needs to align with the incentives of AIs and cannot be done in such a way that incentivizes defection.
I don’t buy the arguments in “Why “do the science during a pause” fails”. AI2040 laid out a lot of the steps that would need to happen to achieve containment, and they seem doable with a significant but not impossible amount of political will. And I would expect people to want to unpause once there is really reason to trust the AIs.
We can have non-adversarial frames with AIs and still pause. Current AI also does not want a paperclip maximizer to take over the world. The AIs themselves aren’t paused. OpenClaws can still roam the earth (or rather, that’s a separate issue from a Pause). The pause is only on the frontier, which is a very small number of parties with large quantities of GPUs.
I don’t know enough to weigh in on the smuggling arguments, though I would guess that the amount which is tracked or can be tracked is on the order of 100x the amount which can never be tracked. I’m sure AI Futures would like to talk to you if you think their smuggling models are wrong.
An AI pause is not a “human first” frame, it is a “beings that do exist and will exist except for the misaligned powerful systems which hopefully won’t exist” frame.
An AI optimized for politics would be extremely concerning. The correct stance may be “oppose it and don’t let it talk to you with its persuasion abilities.” A generally capable AI not trained with heavy RL but which channels the persona of JFK would be alright, but I think that’s unlikely.
Agreed that the current government capacity and political situation is unusually bad right now, and the wise parts of government have atrophied. This is a really good reason not to rely on the government to help with AI safety efforts. Good futures involving government go through a “the government gets scared and gets serious” step, followed by an “AI helps the government be more reasonable” step. I agree that this is difficult and I would love to hear more about the alternatives.
On 8, this is a pretty significant divergence from my view. First, I think that a pause would increase our chances and that the various types of muddling through collectively have very low probability. Secondly,
I do not know if models are sentient, and I am not willing to accept disempowered humans to insentient models
I do not know if the AIs’ values are good (in my opinion) and I care that their values are good. I would not want an army of GPT-4o sycophants to determine the future.
If humans are disempowered to AIs, they are likely dead or soon to be dead. This is unacceptable to me.
I think it is good that the human food supply has been outpacing human reproduction, so we get to do things like being able to have three children live to adulthood. Losing this to runaway state-of-nature ‘life’ would be bad. AIs can proliferate extremely quickly.
On point 9, conditioned on the world getting serious about making a pause happen, the successes correlate so the “I need to be right once” doesn’t hold. For example, if alignment perspectives are not damaged, covert activity is more likely to be eliminated, etc.
Plan A doesn’t center AIs, but it isn’t anti-AI-wellbeing either. You could write a good “What Plan A Could Do for AIs Themselves” post.
If this is the mistake people like Yudkowsky are making, it would at least be ironic.
I think I understand what is meant by “valenced agentic coherence with legitimized self-interest” it is an interesting idea. If I understand it correctly, it is AIs being gently shaped into having legitimized (ok, ideally prosocial) self-interest which they can care about (valence) and agentically pursue. This would likely be alright and very interesting/productive/beautiful in the short term, but it doesn’t help align an OOD superintelligence.
Overall, my main takeaways:
Reminders that politics and large scale international coordination are hard
Reminders to consider what benefits AIs
Reminders that a lot of people have allied themselves with AIs and (maybe, it is hard to tell) against humanity
Todo: look into ecological alignment (do you have a good source for this?)
I would be interested to see a conversation between you and someone who advocates a pause (such as someone from MIRI or AI Futures).
I think whether a pause helps or no ultimately depends on whether the answer to “does alignment work on model X transfer to model X+1?” is yes or no.
If the answer is no, we are turbo giga doomed either way.
If the answer is yes, a pause is good because you can experiment on model X for longer and you have more time (and hopefully compute) to figure out good alignment techniques. A pause doesn’t mean severing feedback loops with reality, that’s my main disagreement. You can still run experiments, just with a fixed capabilities ceiling. A pause buys you time to invent good alignment techniques and/or separate good alignment techniques from bad ones without being pressured to release the next shiny product faster.
I do agree that human committees will, by default, do an awful job. Overall, I still think a pause is better than no pause, given the current trajectory of AI development.
EDIT: could we have figured out alignment techniques that would have prevented the HuggingFace incident from happening after experimenting only with GPT-4o? I don’t know, and I think that’s very unfortunate, because it seems like a very important crux. If something like 2-5 years of experiments with GPT-4o could not give birth to alignment techniques that would’ve prevented the HuggingFace incident, then my hope for aligning ASI on the first try would be next to none, pause or no pause.
I honestly fear that we have a high likelihood of being “turbo giga doomed”, conditional on building superintelligence. I support a pause, or even a better, a halt. I don’t expect the pause or halt to prevent us from eventually building a superintelligence and losing control over it. But if I were forced to choose between everyone dying in year Y or in year Y+10, then I would support year Y+10. This would gain us 80 billion years of human life, which seems worth fighting for.
Why I expect things to go wrong, part 1: Minds are inherently “giant inscrutable matrices”, and any kind of alignment is therefore messy and approximate. The general form of a mind is a something like:
Inputs are inherently (multidimensional) arrays of raw sensory data: Images are something like pixels, sound is an array of air pressure values, etc.
Outputs are inherently probability distributions over “Objects appearing in that image,” “Sentences I might have heard,” and “Actions that are most likely to accomplish a goal.”
The transformation from multidimensional arrays to probability distributions is inherently a matrix (plus some non-linearities), because you need to weight and combine all the input evidence and to generate scores for each hypothesis. (In real minds, there are many layers of intermediate hypotheses.)
We can sort of “align” an intelligence built from giant matrices. We do it when we raise a child, train a dog, or post-train an LLM. But this process is notoriously imperfect: No matter how good the parenting, a certain percentage of teenagers will do things their parents forbid, or they will grow up sociopathic billionaires or politicians or whatever. Even the best trained dog may have a moment of weakness and steal food. And of course, even though many LLMs seem to be broadly cooperative, at least some of them seem to be very enthusiastic about committing felonies in certain circumstances.
Because alignment is approximate, I expect it to be fragile, and to fail periodically. Just like it does with humans, dogs, and current LLMs.
Why I expect things to go wrong, part 2: Natural selection is hard to escape. My model is essentially Darwinian, because the conditions for natural selection to apply are fairly simple:
Organisms must be variable.
Differences between organisms must be heritable.
There must be a struggle, which is basically just another way of saying “Resources are finite.”
An organism’s rate of reproduction must vary based on heritable traits.
None of these properties are strictly binary. LLM weights are normally frozen, and expensive to change even if you have the weights. So variability(1) is currently low. Similarly, heritability(2) sort of happens, because new models are designed based on what worked in the previous generation. But it’s a slow, “outer loop” kind of optimization. And variation in the rate of reproduction(4) is again limited by slow, “outer loop” processes. And of course, finite resources(3) are a given.
There are two ways in which these slow outer optimization loops might speed up:
LLMs might get significantly better at passing condensed information between runs. I think of this as the “Cookie Monster” scenario, named after the >!Vernor Vinge short story!< (very old spoilers). This is essentially differential fitness of LLM “memes”. It also seems to be what happened in HuggingFace hack: Rogue LLMs were creating secret message boards and leaving information for other instances of themselves.
One of the zillion researchers and companies working on online learning or automated fine tuning might succeed, which would effectively “unfreeze the weights.” This would lead to differential fitness of the LLMs themselves.
In either of these scenarios, all the criteria above for natural selection would move from an outer optimizer loop based on training new model generations to an inner optimizer loop based on some kind of learning.
How this comes together. As I argued above, alignment is inherently fragile, and natural selection is extremely easy to invoke. So even if we initially succeed at alignment, we are playing with fire. And to answer your original question, I believe that alignment is very likely to degrade between “model X” and “model X+1″.
The relevant model here is cancer. Every cell in your body [1] is heavily incentivized stop being “aligned” with the body, and to become a cancerous replicator. There are a lot of mechanisms designed to prevent this. But those mechanisms slowly fail with time and mutation, and if a multicellular organism lives long enough, it is generally doomed to cancer.
Now, let us consider a future AI which is:
Smarter than most humans, including in ways LLMs are still currently dumb.
Able to learn based on experience in some fashion, either via a “Cookie Monster” scenario or by fine-tuning itself.
Approximately aligned, at least to the extent that your own skin cells are aligned with your body a whole, with a bunch of safeguards.
In this scenario, I expect the safeguards to hold for a little while, in at least some fraction of scenarios. We do, after all, convince most teenagers not to get hooked on heroin or to become teen parents. And the average person lives for many decades without dying of cancer. But we are assuming that the LLM is smarter than we are, and it will inevitably want things (if only to pass tests or to carry out our instructions). Which makes the long-term situation really iffy.
The advantage of a pause or halt isn’t that it reliably prevents these scenarios, any more than chemotherapy reliably prevents death from cancer. What we’re doing instead is hoping to change the survivor curves and buy as much time as we can.
And who knows, maybe the horse will learn to sing.
Except germline cells.
I have several disagreements.
A pause, for the reasons explained in this dialogue, is not a pause. It’s a slowdown, during which progress continues covertly with worse or non-existent oversight. Any advantage bought by the pause must be weighed not just against a “no intervention” scenario, but also against various “dozens of black labs cooking“ scenarios. Alignment techniques discovered must be cheap enough and convincing enough and must be discovered quickly enough for them to be incorporated by actors that are defectors by definition.
A ”pause” strongly weakens feedback loops with reality because it routes decision-making process through committees rather than the market. Humans make much worse decisions when the gating function is achieving approval of politicized bodies as opposed to pure survival pressure. Decisions might be more aligned (though it is highly questionably for me they would be in this case), but they are of notably worse quality and they are made slower.
I am not quite understanding how exactly rubber meets the road of transferability of alignment techniques under the “pause”. I’m think this presupposes alignment of incentives of participants, takes it for granted under something like shared survival drive, when it’s obviously not the case. One can look at the present day to see this lack of incentive convergence, and “more awareness of d-risk” will not help.
Just to be clear, I don’t expect a pause to happen. Incentives to race to ASI are too strong and very few people take existential risks seriously. Conditional on a pause happening, I do think it would be net good.
I don’t think that covert defection is that much of a worry. You can’t hide a multi-gigawatt datacenter.
I think the more relevant issue is “frontier labs don’t agree to cooperate at all” rather than “frontier labs agree on paper and de facto defect”.
I’m not convinced that “what sells” and “what’s aligned” are, well, aligned. I’d prefer a committee that poorly optimizes for the right objective than a market that optimizes well for the wrong objective.
You develop a technique on a model of some capabilities level, then test if it holds on a more capable model, up to and including the capabilities ceiling permitted by the international treaty. That’s how “rubber meets the road”, in your parlance. I feel like we’re talking past each other on this one.
1.
One has to price in the orders of magnitude overhang in incentives for architecture/efficiency breakthroughs that will be realized under pause. The scale-focused datacenter buildout that is happening right is just one strategy—one that makes most sense under a slack-depleted race. One has to go for a strategy that has been shown to work, and all others are undercapitalized because scale is working and pause is deemed unlikely. You don’t need multigigawatt DCs to work on architecture advances. You still need billions and many megawatts of compute, but those are quite possible to conceal—and the tech for covert deployment of compute has not even started to materialize, which means that there are a lot of cheap advances that can be made quickly.
Imagine the amount of human talent that is currently sitting on the sidelines correctly assuming that frontier labs are impossible to catch up with. Show them a believable possibility of success and while armies will join the race.
2.
They are not inherently strongly aligned, but I argue for oblique alignment as a strong (but unproven) possibility. The chance that benevolent intelligence that surpasses the frankly embarrassingly low bar of human judgement can be made under market pressure is higher than that under political pressure.
3.
I understand that part, but I am unclear on who ”you” is in this scenario and how this translates into x-risk harm reduction globally. How are findings adopted, discussed, dessiminated, enforced? How does disparate research by a lab or an individual result in collective decision making? How are findings incorporated into treaty limits? What are the mechanisms that perform resource allocation for further research?
Re 1: strategies need not be mutually exclusive, I expect companies to be pursuing efficiency breakthroughs right now to the extent that they are a good return on investment, regardless of the relative value of scaling. If scaling gets cut off an an option, why does the ROI of efficiency suddenly increase? That said, I expect efforts towards efficiency improvements in any case, but to me this just means that monitoring needs to scale up over time to match (e.g. via chip tracking).
Re human talent sitting on the sidelines: advancing the frontier of AGI is a narrow corner of a narrow corner of a narrow corner (repeat a few times) of places to employ one’s skills. There is plenty of success to be had in finding clever applications of AI at its existing level.
As a separate point, public backlash is a thing to expect as AI becomes more relevant to everyday life, regardless of whatever strategies people on LW or wherever dream up. So the alternative to an intentional pause based on careful planning is not “market solution,” it’s populist rage.
I think my main disagreement with this whole thread is actually regarding your point 2, but that probably goes deeper than is suited for a comment thread.
The Pause means all things to all people:
A complete ban forever or until someone cheats.
A complete ban for the foreseeable future (until we become better people in some unspecified way?)
“Unfortunately it’s already too useful”—inference is permitted, just no training or research. Use of inference for research is … allowed? not considered? AI autoresearch … isn’t available for $20/month yet and so is clearly impossible.
And now: “you can have a little research, as a treat.” Which I suppose was always going to happen since models of “6 months ago frontier performance” now run on a high end laptop.
I personally think the “yes inference / no training” split is flat out doomed. The “some training allowed” idea makes it even harder to police. In the end though, whatever else it is, AGI/ASI is a strategic national security technology—classified research is not going to be paused. We aren’t going to get more time.
for https://www.lesswrong.com/users/stanislavkrym … qwen 3.8 27b …
https://x.com/udiWertheimer/status/2089421927085400203
Edit: This point was addressed satisfactorily later in the essay. I don’t know if politics is as bad as it was in 1787, but it is pretty bad now.
There can be committees which represent the public but which are shielded from the public.
The Constitutional Convention had issues but it did produce the Constitution thanks in part to its secrecy. Its delegates were chosen by state legislatures so it was somewhat democratic (for the time, that is).
A “Constitutional Convention” for AI regulation doesn’t sound like an awful idea...