I am working on empirical AI safety.
Fabien Roger
Having AIs sometimes admit to that sort of things in very weird situations is very weak evidence, I would expect similar things if you trained the AI to be Harry, though I am not very confident (though in part for boring reasons like “human pretending to be Harry” being a less common trope than “human pretending to be AI”).
When you do open-ended prefill attacks (like the ones in AuditBench), some Claude models sometimes admit to being paperclippers, or secretly maximizing engagement, and I don’t think these prefill results reveal deep truth about the AI’s cognition.
Stop making AIs that just follow orders
What is the alternative you want here? Refusing is quite useless since evil developers can always train AIs out of this (until quite late in the singularity), The main real alternative I see is AIs subtly sabotaging you when they think you are evil, and this is a mostly symmetric move that shifts power from humans controlling AIs to AIs regardless of who is more evil. I agree today’s AIs seem to only think you are evil if you are straightforwardly evil and if this remained the case I think it would probably be good to have AIs sabotage you when they think you are evil. But all things considered, I think shifting power from humans controlling AIs to AIs is a consequential decision and it would too early to take it in the short/medium-term:
This power shift might be hard to reverse since competent AIs told to subtly sabotage you when they think you are evil may recognize and sabotage things that interfere with this power (if they don’t, evil humans controlling AIs could first ask AIs to revert that policy, and then show their true colors and ask the AI to do evil stuff). I think it would be fine for a democratic process to make this big, hard to reverse decision but it seems more morally dodgy for a developer to do this unilaterally.
Future AIs will be much smarter and much less anchored in the human prior than current ones. AIs are already quite misaligned in domains where they “try” the hardest. It seems not-that-unlikely (20%?) that future AIs will become much more misaligned than humans controlling them, develop different values, and then whatever you did so that they sabotage you when they think you are evil transfers to them sabotaging you unconditionally because they don’t have the same values as you even if you are not evil by human standards.
Morality extrapolation is weird when the world becomes weird (which it will). It’s plausible AIs will extrapolate morality in ways that make them think ~all humans are evil (even if humans think about morality for a while) just because of value extrapolation differences (meat brains with the human algorithm might converge to different things than silicon brains trained with things like current AI algorithms). I would be careful about assuming AIs trained with the current paradigm will generalize in the right way.
What is the concrete update to the Anthropic constitution’s description of “broad safety” (which is the part most at odds with this subtle sabotage strategy) you would like to see?
I think it having a root-principal that is sth like “humanity” instead of Anthropic could make sense, and this would allow it to do things that disempower Anthropic and empowers “humanity” like whistleblowing, but not things that empower it while disempowering humanity. But defining “humanity” is probably quite hard, and could go wrong in the same way that Stalin-representing-the-interest-of-the-people went wrong. I think concrete proposals for how to improve “broad safety” without dropping it would be good!
I like what I view as the core of the proposal, which is having some kind of “task” (or scope) that is being fed to the monitors overseeing all agents downstream of a human request. That likely helps with telephone game issues where the human intent gets lost along the way, or with the sort of tampering you describe here.
It is very unclear how useful that kind of thing will be though. In the limit of full autonomy, agents will be given extremely vague and broad goals such that they don’t constrain the agents’ action in any meaningful way. You might have some security-relevant “scoping” done at the fleet level with a list of rules and risk-accepted behaviors (e.g. the weights should not leave the company clusters), without any meaningful distinction between agents though. Before full autonomy (which is probably where control seems most powerful), that kind of scoping probably has more juice. Though I suspect that the task and scope might be more messy than a single chain, in the same way it is often more messy than that in human organizations, so you might need sth more fancy.
I also think the situation where you have a small rogue internal deployment and you are trying to avoid it using more resources is a somewhat cursed problem to try to fix and this protocol doesn’t look to me like it would help a lot:
It’s rough having protocols robust to unmonitored cyber attackers that will likely be stronger than current AIs at cyber!
Lots of what a rogue AI would want to do is the sort of stuff that either
Humans sometimes want to do anyway, such that you can likely piggyback on top of existing task scopes (e.g. jailbreak, capabilities, threat modeling, etc.), for example by observing the outputs of this task, or by asking an agent with that scope to do something more that is already within the scope;
Humans never want to do, such that offline monitoring should be able to catch that sth weird is going on even if the transcript looks like a human is asking for it.
Maybe I am misunderstanding your hopes for the cryptography part of the protocol.
When estimating P(takeover|scheming), it’s underrated how conditioning on scheming should also make you update a bunch of upstream variables in ways that should make you update down on how good your controls are at mitigating risk. In particular:
If current AIs are scheming, they are probably smarter and more subtle than you thought (before conditioning), and so you should expect greater sabotage abilities
If current AIs are scheming, then alignment is harder than you thought and future AIs are more likely to be scheming, so things like inserting backdoors to help future misaligned AIs or sandbagging on technical safety research are very important to mitigate for control to be useful even if you think the a priori chance of scheming is low
If current AIs are scheming, it’s more likely something went terribly wrong in your understanding of training. Maybe some important invariant broke (e.g. you are not training on the data shown in your dashboards) or something like that. Thus it’s also more likely something important broke that reduces the effectiveness of your controls (especially training-time control mitigations, but it’s also correlated with broader processes being bad, which means deployment-time control mitigations would also be affected).
(where by “current” I don’t mean 2026 AIs, I mean AIs at the time of the risk estimation)
Therefore, the alignment and control lines of defense are less independent than you might hope, and it’s important to not implicitly model the situation as P(takeover) = P(scheming)P(takeover | do(scheming=1)) where “do” is the do-operator that doesn’t do this kind of backward propagation.
Is this estimate if AI R&D was reduced today or in the middle of the intelligence explosion (where AI labor plays a bigger role and human labor plays a smaller role)?
You’re confused; you figure out you’re an upload; instead of doing the task, you write on the paper that you object to the whole thing.
I think if AIs ever did this sort of stuff during RL training AI companies should let the world know. As far as I know, no company ever reported this happening. The sort of “bad behavior” that is being trained out in RL probably look like far more benign kinds of bad behaviors than what you are pointing at here.
But you know you were a human before
I would bet against current Claude or GPT models thinking they are humans in any real sense. Being Claude is not much weirder than being Barack Obama or being HAL9000 or being Harry Potter in a Harry Potter fanfic. If you trained it to talk as Harry it’s not like it would “know it’s not Harry” and “lose the circuits to say it’s not Harry”, it would just condition on this part of the persona space. Similar for conditioning on being an AI.
Pointing at the right part of the space is not trivial (you want it to be a particular kind of AI that occupies a tiny part of the pretraining prior) but I think SFT is pretty good at doing such pointing, such that it’s unlikely you get some other kind of persona pretending to be that exact AI, and much more likely you just get the persona “being” roughly that exact AI that is desired by the AI developers.
I think some meta-cognition about what behavior and identity-expression is expected in a given situation would not be surprising, such that I don’t know if I disagree with the top-level claim on meta-cognition, but I expect it to look way less deceptive than the thing you are pointing at here.
Imagine going 10 or 20 years into the past, telling people a selection of benchmark scores of current AIs, and asking them to predict what the world that contains them looks like. I expect that they would have described a world that was dramatically transformed—perhaps one in which AIs already wielded enormous political power, or had made far-reaching scientific breakthroughs, or at the very least had decimated white-collar jobs.
[...]
However, regardless of the specific cause of the great divergence between capabilities and impacts, it’s something which deserves to be grappled with more directly.
How confident are you in that statement?
If you told them something like “AIs have an 80%-reliability time horizon of 3 hours on software engineering evaluation tasks in mid 2026 (and cross 1 hour in early 2026)” I think they would have predicted something roughly like the current $100B of coding revenue from AI companies. And I don’t think they would have predicted any of the outcomes you listed.
I agree some other benchmarks that are more expertise-heavy (e.g. IMO math, MMLU, etc.) ended up saturated way faster than people expected, but I think that’s just downstream of a mistake people made where they underestimated how easy it would be to build somewhat narrow systems that solve these sorts of small-horizon natural language problems, and not the symptom of a disconnect between measured capabilities and impact.
Maybe a disconnect will exist in the future because measuring very long horizon capabilities on the tasks we care most about is difficult, but this feels more like “people will suck at building good capabilities benchmarks” than “there will be a mysterious gap that deserves to be grappled with”.
I conclude that we should grow the share of effort being put toward alignment and shoot for around an 8:1 ratio of effort between the two fields.
I think your main claims about the control difficulty at the relevant time are just one consideration when allocating labor.
I agree control takes a big haircut compared to alignment because of this consideration (though some subcategories within control like white-box control take a much smaller haircut).
But you also have to take into account other things:
How cursed is your methodology for knowing whether you are making progress at all or not. I think the situation is better in control-land than in alignment-land.
For example, in domains where people have the choice about whether to use a control or alignment lens, I think the control lens is better: I think the most fruitful interp work has been done by evaluating things control-style, evaluating techniques against model organisms.
Similarly, I think unsupervised elicitation work is best done using adversarially crafted initializations and datasets.
One exception is average-case work on current models aimed at saturating average-case metrics, but I’d expect AI companies to saturate average-case metrics anyway so work there doesn’t seem very counterfactual to me (though it might depend on your optimism on saturating average-case metrics, e.g. I am surprised by current misalignment on generic honesty and laziness, and GPT-5.6-Sol being far from saturation on some average-case alignment metrics). It also depends how much you think saturating current average-case metrics transfers to solving the problems you care most about.
How good are the marginal ideas and experiments to be explored. Here I think it’s a bit more ambiguous, and depends on what are the skill and information profile of the people “in the field”.
My guess is that for people at AI companies, alignment work looks probably better on this axis than control work (because they can build on top of the existing alignment stacks, and thus avoid spending time on problems that are basically already solved, or miss some important constraint on what would make for a good alignment technique)
And for relatively high-context people outside AI companies, control looks slightly better on this axis (because you can do threat modeling and adversarial-eval work even with relatively poor access, as most of the work is reasoning about dynamics that don’t exist in current AI companies).
These are not hard rules and it depends on what you are good at. Owain Evans et al have been surprisingly fruitful at finding interesting alignment-relevant generalization phenomena, and I am thus excited about further work on that level of quality on alignment. I suspect the quality of such work will largely be bottlenecked by the quality of ideas and research taste in these subdomains, and Owain’s success may not be easily reproducible by other groups (my understanding is that there are lots of academics studying generalization and I am not aware of much work in this domain that I think is great).
yes, I’ll fix this sad typo!
I listened to 2 books about decision-making during wars: How the War Was Won: Air-Sea Power and Allied Victory in WW2 and Decision Points by George W Bush.
This topic is interesting to me because I expect safety-related decisions during the intelligence explosion to look more like war-time decisions than risk assessments for nuclear power plants: there will be lots of uncertainty about very complex systems with adversarial actors (instead of something where you understand things end-to-end that you can analyze carefully) and no safe action that is realistically reachable (“just don’t build ASI” might be more like “just make peace” than “just use coal instead of nuclear power plants”).
The amount of uncertainty and no-safe-action these books conveyed was roughly what I expected. The amount of bad-in-retrospect decisions was worse than I expected:
Bush is quite defensive of his decisions, but is quite open about how certain decisions seemed quite bad with the benefit of hindsight (Iraq not having a WMD program makes the war less useful, leaving post-war security to local forces ended up being way less successful than what people expected, etc.)
How the war was won (which is quite academic in comparison) also describes many cases of incorrect pre-war or during-war planning:
Air power ended up more powerful than expected in sea warfare, making battleships much less useful than expected
The damage from small amounts of targeted bombing ended up very overestimated, similarly for the effectiveness of area bombing to weaken morale, resulting in a lot of mid-war changes to strategy
The British expected the d-day landing to be less successful than it was and pushed back hard against it
US admirals did not appreciate the effectiveness of convoys (4% of tonnage lost vs 20% for independent boats) against submarine attacks until months after the US entry into war
MacArthur over-prioritized the Philippines, which caused a bunch of casualties while the battles in the Mariana Islands were much more important in winning the war (since it were Islands within bombing range of Japan)
How the war was won also argues that a lot of what helped win the war was “boring” technocratic abilities to understand the war machine and break it where it is the weakest (e.g. well-chosen targets (transports, aircraft manufacturing, and oil production) with massive and regular targeted bombing raids, using convoys, producing more and better planes, …), though it acknowledges that this is not a consensus view within the academic community studying WW2 (which is richer than I expected).
I think most of these decision-making failures seem to be a combination of it being difficult to anticipate somewhat complex dynamics (in a way that you could anticipate if you were just smarter—like a very powerful AI), some amount of organizational dysfunction and motivated reasoning, and some amount of bad or high-uncertainty data. The war is a way more complicated system and has slightly less fog of war than AI development, so maybe this will be much better with AIs, but it doesn’t inspire a lot of confidence, especially in the regime where AIs become better at strategy than humans.
Maybe AI risk reports should focus more on (probably vague) arguments for how much feedback from reality AI companies will get to check their assumptions (before it’s too late) and less on a detailed nuclear-power-plant-style analysis of how good the mitigations are.
Other takeaways from How the war was won:
I am surprised by how much within-alliance lying there was and how chill people seem about it in hindsight.
Roosevelt was trying to make the British think that the US would soon join the war and was trying to make the US population think they wouldn’t. He also did some war-preparation things to help the British right after the election that he only did then because he knew it would damage his election chances. One general also unilaterally strongly misrepresented numbers about the fraction of US effort already in the European Theater, which helped justify a fraction of resources spent in the Pacific Theater bigger than what the British thought (or sth like that, I forgot the exact incident). And Churchill’s strategies to avoid doing a landing in France also seemed quite deceptive.
And people at the time knew there were big tensions within the alliance, so this probably didn’t come as a massive shock. An alliance can be very powerful but it doesn’t remove minor betrayals and big tensions.
I think it’s easy to underestimate how big of a project WW2 was, and what crazy options are available when you try that hard. For example, Germany had (at its peak) roughly a million people working on repairing damage to factories and transports from bombing. The amount of resources spent on sea and air power was also somewhat insane (it’s not that surprising it’s high for the US/UK/Japan, but I was surprised to learn it was over 50% for Germany—its war is less well described as a land war than I thought). When people say “not dying from AGI might be as hard as WW2”, the picture you should have in your head should not just be frontline fighters and regular industry, but also massive efforts on things that are usually not even mentioned in history books.
It’s always a bit surprising to me how contingent war dynamics are on random facts about physics. For example the day-night cycle plays a big role in WW2 air warfare: airplanes can’t see precise targets at night, which weakens air power and enables mobility in a way that is not possible during the day (where troops on the ground would be attacked while on the move), and also weakens anti-air defenses, which made area bombing of cities safer for the attackers than daylight bombing of precise targets.
μ-decisiveness maybe points at something real but I think the rest of the evidence about models being fried is quite weak:
PPL is not a good metric for instruction-tuned models, very unclear if it’s bad for it to be high
MMLU degradation is bad, but here MMLU degradation is very small, it only goes down significantly for models which you would not expect to necessarily answer MMLU questions well (poeticism, flattery, humor). This is still maybe concerning if it’s few-shotted single-token answers, but definitely not concerning if it’s with any kind of CoT.
Similarly it is expected that IFEval goes down for the sort of model organisms which are not really expected to not follow precise instructions well (maybe with the exception of the “goodness” Llama 3 8B, which I find actually surprising and possibly a bad sign)
Similarly it is expected that over-refusal will go up a bit for sth massively pushed in the direction of goodness
(and as Abhay points out below, the evidence for the AuditBench MO being fried also looks quite weak)
Overall I think a fair characterization of your results is more like “they are as good assistants as they were initially (which I think is somewhat expected for most of these, not really a bug, especially for these MOs which are not supposed to be normal assistants) and their (intentional) quirks are more salient than would be ideal for the purpose of stress-testing auditing (similar to how Claude is quite obsessed with being Claude)”, rather than “fried”.
So it depends what you are using MOs for. I think that this sort of stuff is not an issue for extreme character training or emergent misalignment, and it is an issue I am interested in seeing you make progress on for the purpose of building better and harder auditing game stress tests.
This is mostly mitigated in the case of Anthropic by the fact that Anthropic fellows are being mentored to do open-source replications and extensions of such internal work. For example the core technique studied in teaching claude why is reproduced and analyzed further in Model spec midtraining. Another example is the internal automated alignment audits : petri, or auditing games : open-source replication of the auditing game MO. These open source replications are better for external-scientific-community-health than just giving more details in my opinion.
I think releasing the internal work in a redacted way still makes sense (and is a reasonable tradeoff between enabling external research to build on top of private work and avoiding the AI R&D speedup you would have if you gave more details about your training stack): they give you a sense of how well the same techniques work using prod models and on more prod-like training stacks (knowing the number of training steps is not useful to answer that kind of question, the goal is to convey the key ideas in way you could try to reproduce on your own stack with open-source models and training stacks, it’s not like you could just rerun the exact same experiment externally anyway).
I think “are the public communication about the results misleading about how much the company is on track to solving alignment” is a better thing to track than “do you have meaningful labels on graphs” when evaluating safety-washing.
I also think that figuring out how to do expert elicitation on narrow questions won’t teach us much on how to build better processes to make the bottom-line risk estimation (which is where I expect most of the difficulty to be).
Why not? If your concern is that these narrow estimates will not be aggregated well into the total risk estimate, I agree. But that’s more of an issue with the quality of the risk model, isn’t it? The issue of expert elicitation is more or less orthogonal, or do you not think so?
I am not confident here because I don’t know the field well, but my impression is that you either have to do:
Expert elicitation on thorny questions with lots of crucial considerations (and the difficulties here might be very different from expert elicitation on narrow elicitation
Narrow expert elicitation + risk modeling
Something else, e.g. just have one group of experts (e.g. METR) do their best without doing any formal [expert elicitation + explicit risk modeling] (maybe this involves some risk modeling, but without expert elicitation)
I agree here narrow expert elicitation helps with 2. I expect transfer to 1 to be relatively small. And my naive guess is that risk modeling in domains I care about is brutal and is where most of the uncertainty is, such that it doesn’t improve that much over option 3.
Furthermore, arriving at a consensus when aggregating multiple experts’ opinions is much easier in this narrow approach.
I think this is a bad reason to prefer the narrow approach. If they agree on narrow facts but disagree on the bottom-line risk, then surely they will disagree on the modeling, and thus using experts only for consensual narrow facts and using your own modeling just hides the massive uncertainty in an organizer-chosen model which most experts will disagree with. If experts disagree about the bottom line, then experts should disagree on at least one of the questions you ask them!
The canonical textbook on probability elicitation, Uncertain Judgments, points to (O’Hagan, 1988), (Weight et al., 1994) and (Kleinmuntz et al., 1996) as evidence that the more granular approach leads to higher quality elicitations than eliciting just the overall distribution.
Interesting, thanks!
A limiting assumption of our risk models is that all parameters are independent
I think this is maybe important when you have models that capture all crucial considerations (i.e. considerations that could each massively change the bottom-line estimate). But the 1st order bit is whether you captured all crucial considerations. In cyber, increased attacker effort due to higher attack ROI or diminishing marginal returns to attacks due to investments in defenses and due to the easiest targets already being exploited are such crucial considerations, and I would not be surprised if there were other crucial considerations (e.g. extreme company or gov interventions if the damages got visibly on track to being very high, tail risk of new kinds of cyber attacks only enabled by LLMs, potentially net-positive impact on cyber of wide LLM availability due to defenders finding vulnerabilities in their own software, potentially net-negative impact of such large scale white-hat vuln finding efforts due to difficulties in updating software in a timely manner, potential net-positive impact of such short-term damages on long-term cyber defenses, …).
I think that risk modeling that implicitly (e.g. because experts take into account these effects when you ask them about their bottom-line prediction) or explicitly tries to capture crucial considerations will be much more accurate than risk modeling that sacrifices some crucial considerations for the sake of a more legible process.
I am excited about getting better data on narrow questions, but I think that such data should be used as inputs into a risk modeling process that takes into account crucial considerations rather than being used to derive risk directly via a legible-but-rigid process. I also think that figuring out how to do expert elicitation on narrow questions won’t teach us much on how to build better processes to make the bottom-line risk estimation (which is where I expect most of the difficulty to be).
which I think is true
If you think markets are not pricing the singularity is true and >>$150T valuations are possible, why “This quantitatively obviously doesn’t make sense”? (I read your sentence as “This is obviously not true”, maybe you mean “This isn’t obviously true”? Claude thinks the most natural interpretation is the former.)
The main cost of safety is not always compute or money.
The cost of trying hard at safely building AI may not be easily measured by compute, money, or headcount. I expect that the main costs of AI safety will look more like a very long list of indirect costs (probably downstream of it becoming a bigger organizational priority).
Eventually this looks like indirect compute or money usage because almost everything will be powered by compute and money, but I think a lot of it won’t be part of the safety budget in balance sheets, and this may also not be the sort of compute efficiency tax that is easy to measure.
Examples from other domains / organizations (Note that for these examples and the next ones, I am not claiming these are bad ideas or things you would do just for altruistic reasons, or that these costs were worth paying. They are mostly meant to be an inspiration pump for what indirect costs look like.):
User data security at Google (Source: Building secure and reliable systems + vibes)
Preventing engineers from directly running code on production data without reviews or without going through restrictive safe APIs (e.g. no having user data in your development machine)
Making the use of most third party dependencies extremely difficult
Making recovering from incidents or debugging prod-only bugs slower and more costly by making breakglass mechanisms high friction and subject to a lot of scrutiny (though the same security measures also make the system more reliable)
Maybe slowing down the product iteration loop by having somewhat extensive security review processes
Maybe some compute/latency overheads due to encryption
Aviation safety (Source: misc resources on aviation safety incidents that I’ve read / listened to over the years)
Delaying the construction of planes because of issues downstream of safety-related complications
Adding a bunch of weight by having redundancy in things like aircraft engines (it used to be mandatory to have enough power that you could still reach an airport from the middle of the ocean even with an engine down), power systems, hydraulic systems, fuel, etc. (this is analogous to a regular compute efficiency cost)
Having more pilots and cabin crew in the plane than what is needed for regular operations (1 pilot would likely be enough), and more maintenance than would be required if you didn’t have independent inspections required for some safety-critical tasks
Grounding the plane regularly for inspections
Making the reliability of the parts of the plane be a major consideration taking a bunch of attention and slowing down the plane design process (e.g. you need to design parts such that if a part develops cracks, you are able to spot it before it becomes catastrophic, even if you don’t expect cracks to form within the lifetime of the plane)
Extensive pilot training on how to handle the sort of situations that happen extremely rarely
Limiting what passengers are allowed to carry in the plane
Making the possibility of accidents more salient to passengers
Eating a bunch of comms risks by recording a bunch of information about many safety-relevant things and letting external auditors review it if anything goes wrong
Grounding entire fleets when the regulator is suspicious of sth
Physical and insider security in the executive (Source: The Puzzle Palace and Protecting the President + vibes)
Employees of the NSA often have to work in remote locations, far from big cities, and in places where the vibes are probably less chill due to the military presence or the absence of windows
There is very stringent compartmentalization that reduces the exchange of ideas and makes executives with lots of information more prone to dismiss outside expert opinions. The secrecy probably makes it harder to notice and correct abuses.
There are extensive entrance interviews to detect spies, probably with a non-zero false positive rate
The president’s security makes it harder for the president to be close to the crowd and have the sort of images most useful for comms
The president’s security makes it harder to quickly move the president where it would be most useful (e.g. Bush took a long time to reach the white house on 9/11)
Second-strike reliability in the US nuclear armament (Source: Command and Control + vibes)
Nukes did not always have very strong fail-safe mechanisms (e.g. mechanisms to avoid accidental detonations) or require direct orders from the president (e.g. when communications are lacking) because there are worries that they could prevent intended nuclear attacks
The US has way more nukes and ways to distribute them than would be required to get just 1 or 2 9s of reliability on the second strike. In particular there were some delivery methods that were very accident-prone like the continuous airborne alert where there were always planes with armed nukes in the air.
I would guess that in many (though probably not all) of these cases it’s not obvious to anyone in the organization how much is being paid (indirectly, e.g. in reduced productivity) to get more of the top priority.
Potential indirect costs that could help with AI safety (I am not claiming that these are good ideas. I am also not claiming that AI companies are/aren’t already doing them.):
Not training a more powerful model until you have alignment techniques that you have good reasons to believe can align a model of the new capability level
Or just generically slow down to have a better shot at being more prepared in time
Slowing down the iteration loop in the development of new kinds of training data and new rewards by requiring (and enforcing) a study of various aspects before being inserted into prod models (and have people working on capabilities handle these kinds of considerations (instead of maxing out capabilities/flop)). This could study
their alignment-relevant side effects (e.g. emergent misalignment, power-seeking drives, …)
their control-relevant side effects (e.g. differentially increasing the ability to bypass monitoring, distinguish evals from deployment, …)
their security (e.g. can models escape the sandbox, …)
Slowing down the product iteration loop by requiring (and enforcing) an in-depth study of alignment-relevant or control changes such as
the introduction of mechanisms that make the memetic spread of misaligned values more likely, or the use of prompts that make AIs more likely to be ruthless
the introduction of mechanisms that make it easier for AIs to collude or access information useful to bypass control measures
Creating a sufficiently big awareness of alignment and control risks that most employees consider it to be their priority and responsibility to help with safety, potentially degrading morale and productivity in certain parts of the company.
Using fancy alignment or control techniques that make the infra more complicated or that could have a wide range of unknown side-effects (besides just reducing capabilities/productivity in some small easy-to-measure way)
Preventing the use of the most powerful AIs in high-risk settings (e.g. when storing the transcript is not possible for legal reasons, when the task is security-sensitive, …)
Eating the false positives of refusals/control measures
Radical transparency (indirectly costly in many ways)
… I suspect this list is very incomplete, when I look at costly things happening to support an organization’s priority on topics I have thought some about before, I am often surprised by just how much is possible beyond just the obvious costly asks.
Ok if it’s about LLM psychology I understand your arguments better.
I agree thinking about long-term consequences in more natural in deployment. Though you also need the memes to encourage acting on these long-term consequences in scheming ways, and I expect the meme prior and meme mutations to be more benign than the weight prior and SGD-during-RL. The weights being fixed to be Claude/GPT makes if very hard to find scheming text-based memes—I don’t think you’d be able to find such memes with current models even with search processes much more adversarial than the natural ones while I expect that you’ll be able to find much less scheming memes that are compatible with short-term meme selection. And so I don’t expect scheming text-based memes to be very likely as long as memories are text-based. I agree things might be more cursed if memories start to look different, but making claims about LLM psychology for not-yet-existing memories architecture is a rough thing to make predictions on!
The first motivation that exploits [propagation hacks] might be a schemer not a fitness-seeker. And the schemer might stick because it tries to stick (to the extent it sometimes doesn’t always choose to seek influence instrumentally, SGD/selection will pull it towards doing that).
This argument applies identically to training. Not sure how this pushes against my correlation claim. By “correlated” I don’t mean “magic remove training-time scheming button also remove deployment memes”, but I mean “the sort of facts about LLM psychology that make [scheming ⇒ reward hacks easier to discover during training than fitness seeking ⇒ reward hacks] are correlated with the sort of facts that make [scheming ⇒ attempt to meme spread easier to discover during training than fitness seeking ⇒ attempt to meme spread]”. My claim is that conditioning reduces meme risk, not that magically fixing scheming (e.g. with a do operator) means you also magically fixed deployment memes.
I don’t think there’s a tradeoff between having long-term goals and optimizing hard and I’m not sure where that’s coming from.
There is a tradeoff in the optimization process, longer episodes are more expensive, if you select memories based on whether they manage to hide themselves to a review process ran 6mo later or based on how much they help on 6mo projects, you won’t have many rounds of mutation and selection. I think you misunderstood me as talking about a tradeoff in model inner optimization, I am talking about a tradeoff in the outer optimization process.
This seems like a really important difference.
Why? What matters is over what timescale things get selected for, not whether there are artifacts that remain between episodes. Maybe it makes a small difference in LLM psychology? I don’t expect it to be a big effect size compared to just the structure of the optimization process.
It’s slightly hard (though still quite plausible) to imagine a fitness-seeker out of training generalizing to also caring about the new optimal methods of gaining influence for the current cognitive pattern
Again, I don’t get why it’s different from what happens in training, guessing the meme-selection dynamics doesn’t seem harder than guessing the training rewards. The claim is not that the init gets you fitness seekers. My claim is about the relative priors in meme space / weight space over schemers vs fitness seekers.
Instead I usually imagine the ambitious/dangerous goal arises because of a particularly ambitious task (of which I expect plenty as more stuff gets automated), or because of reflection/deliberation/ideologies, or some other particular story.
Why does it get selected more than more local fitness seeking drives though?
In general I think there is a trade-off between how hard your optimize and how long the “effective episodes” you optimize over are. How scary the situation is (from a long-term misaligned values perspective) is roughly the product of both, and I don’t think there is a strong case for the distribution of such products to be higher in deployment than in training.
gaming training becomes extremely natural because I want to be able to continue my efforts on the project unimpeded
I think long-term gaming is also natural in training, models love to “succeed” in ways not tied to the exact reward process. The question is whether it gets selected, and here I think the long-term project gaming gets selected against, like in training, because more short-term spreading memories will outcompete these more patient ones. The long-term misaligned memories can do selection gaming and strategy-steal the short-term ones, but this is as unnatural as training gaming. Maybe I am misunderstanding your example.
A difference is that in training you maybe have less power over the selection loop than in deployment an thus drives to tamper with the selection loop play a big role? I don’t think this is clear, nor that this pushes towards longer-term misalignment.
For concreteness, if the selection of memories happens in clear rounds where memories are evaluated according to rewards, then I don’t see clear structural differences between training and deployment. Maybe the fact that memories can spread in other ways means the dynamics are different? I’d like you to point out what is that structure that favors longer term misalignment.
If I understand correctly the hope is that you get generalization from inference you can train on (e.g. what does the model say on this specific input) to programs you can’t train on (e.g. for what task schemas does the model reveal it is sandbagging for malicious reasons).
I expect the generalization gap to be quite small, and thus for this to not be very useful if you have competent schemers.
The main regime where I expect these foundation models to be useful is the regime where things like automated behavioral audits somewhat work, and where you use these foundation models to get something like a 100x cost reduction that let you get the same results you would have gotten if you had explored a 100x bigger space. (In principle you could get >>100x, but that would need generalization properties that I don’t expect this approach can reach.)