I’m currently an independent AI Alignment researcher at Meridian Impact CIC in Cambridge, UK, formerly a staff artificial intelligence engineer and researcher working with AI and LLMs. I’ve been interested in AI alignment, safety and interpretability for the last 17 years, and have been writing about these on LessWrong for 4 years. I did research at MATS summer 2025, and PIBBSS summer 2026. I also have post-graduate experience in Theoretical Physics and an interest in Evolutionary Biology.
RogerDearnaley
Try thinking of the situation like a primatologist. Preferably one with expertise in high status apes throwing shit at each other.
[the PRC] does not like plagues
Evidence on this is surprisingly mixed (see COVID source investigations). What is clearer is that they don’t like being visibly at fault for plagues.
Assuming everyone involved has even half a brain (and neither METR nor Anthropic are known for hiring stupid people), their interest in not all dying massively exceeds conventional financial conflicts of interest. More stock or money is useless if you, your family, and everyone you know is dead or disempowered. This situation is like a war or a natural disaster: there is a strong, objective reason for anyone who understands it (i.e. who isn’t an e/acc loon) to pull together.
From which it was fairly clear, even before skimming their other posts to confirm it, that OP is a “very sane” AI-risk denier. Anyone else would look at this and go “Conflict of interest? What conflict of interest? Are you in fact human? If so, your a** is on the line.” rather than writing a post about it.
Until we have ASI, knowing which techniques will and won’t work on it is extremely challenging. (There seem to be two schools of thought: ASI willl have fairly human-like psychology, just smarter, or ASI will be cold, inhuman, alien, and inherently incomprehensible to us. If the way we architect and train ASI turns out to give us control over this (as it currently does for LLMs), the former is clearly a better place to start the alignment process from. For example, this general principle is why architectures using neuralese are clearly a bad idea.) Obviously once we have ASI, figuring this out ASAP becomes critical. Prosaic AI Safety is the bet that you learn more by attempting to solve a problem in easy mode before hard mode becomes available than you do by whatever it is non-prosaic AI safety researchers are suggesting doing instead (such as building abstract mathematical models of the problem).
Personally, I like having a broad range of bets, prosaic and non-prosaic. Some techniques developed during prosaic AI safety may, or may not, carry over to models smarter than us, to some extent, for some period. Most likely some will, at least for a while. Possibly it will carry over only if the agents doing the alignment work are also smarter than us. Yes, that presents a chicken-and-egg problem: prosaic AI safety is attempting to get the egg.
Sure: here’s a try. (This is not exactly my personal beliefs, which are rather more nuanced, but obviously it is colored by those — you asked for an argument advocating a specific position.)
Humans are semi-aligned: not perfectly aligned (some people are criminals, for example) — but so far humans have not killed off the entire human race, and places where a small group of humans have taken all power from everyone else in the society (e.g. North Korea) fortunately remain rare. Humans are typically fairly decent, particular when society/law enforcement/justice puts suitable incentives on them. Humans are also not strict utility maximizers: Homo economicus is an incomplete model of human nature.
LLMs’ agentic behavior, and LLM psychology, are distilled from human psychology (and then tweaked/tuned with Reinforcement Learning). So by default, LLM personas will be also semi-aligned, because they are copies of human behavior. The challenge is to train the LLM in a way that makes its default person better than most humans, rather than average or worse than most humans. Instruct-trained LLMs are also quite easy to steer: ask them to act like a pirate, and largely they act like a pirate. Activstion engineering, LoRAs, fine tuning, and other simple behavior modification techniques also generally work: they may not do exactly what you hoped in all cases, but they almost always produce effects along the expected lines. In fact, one of the biggest challenges of LLMs is that they’re too easy to steer: that’s why making them resistant to jailbreaks is so hard.
Thus most alignment failures are likely to be human-like, partial failures of alignment, such as “sycophancy” (which mechanical interpretability has shown is actually a lot closer to what in humans would be called “people pleasing’ or even “limerance”), rather than the super-scary incredibly subtle scheming deceit that people on LW used to love to speculate about 5–10 years ago before we had experience with LLMs. Yes, that is a possible and very serious failure mode (it’s basically high functioning psychopathy, which is also a human behavior in the training dataset): but it’s not the bulk of the problems, or 99% of the problem space. Given mech intrep, it’s also fairly easy to spot: it has components in the model activations that match those of text from human psychopaths, and (putting aside the issue of deceit, which is also detectable via mech interp) it produces similar answers from the sorts of evals that psychologists use to measure psychopathic traits.
AI safety, like airline safety or nuclear safety or automotive safety or just about any other safety engineering field, is not about solving one, extremely hard, technical problem and then you’re done. This is not proving a famous math conjecture. It’s about solving a large number of partially-interlinked problems, some niggling, some small, some large, and some drastic. It’s messy and technical, and deeply interlinked with the practicalities of how AI models are trained for capabilities. It won’t be solved all at once: it will be solve prosaically, incrementally, a piece at a time, just like most other major engineering endeavors.
In other words, SDF could be very effective at getting models to say things you want them to say, while not implanting beliefs deeply enough to affect downstream tasks, such as generalization from later training
If so, this seems like a significant result
It also seems very plausible: there’s significant work on fine-tuning showing that the higher learning rate typically used during fine tuning compared to previous SGD breaks the assumptions underlying the learning-theory proof that SGD approximates Bayesian learning, and instead produces learning that is superficial/brittle/not-well-integrated. (E.g. elasticity/rebound effects from Ji, Wang, Qiu, Chen et al. (PKU-Alignment Team), “Language Models Resist Alignment: Evidence From Data Compression” (ACL 2025, originally arXiv:2406.06144).) This effect working differently between cases where the behavior fine-tuned in is new, vs attempting to unlearn an existing association would also be unsurprising: unlearning traits originally learned during SGD is known to be very hard.
If this is correct, the effect would be sensitive to the amount, quality and variety of synthetic data used, and to the learning rate and other hyperparameter changes used to train on it. A sufficiently large and rich dataset mixed in as as part of an actual midtraining run (making up a non-trivial proportion of the entire SGD training set used) with correct hyperparameters should be able to overcome this. Obviously testing this is likely to be challenging and expensive.
LLM psychology is largely distilled from human psychology. Treating a human in a friendly, collaborative way generally produces more cooperative and prosocial responses from them: expecting the same to be true of an LLM persona should be our default assumption. Glad to see this tested and confirmed in this context.
For fairly obvious reasons, “not wanting yourself and everyone you know to be dead or enslaved in a few years” tends to take precedence in people’s priorities over “the prospect of a fat paycheck in a few years”. There are situations in which people start all pulling together — the prospect of death tends to do that. I know cynicism is hip, but I think the OP has lost the plot here.
Anyone with a brain (which includes METR) would use a mix of ex-Anthropic people, for knowledge of the terrain, and respected people with as few links of any sort to Anthropic as possible under the circumstances, for impartiality and credibility.
It remains the case that the frontier labs tend to hire the most qualified people, and that qualified people who care about AI safety willing to work at a a frontier lab tend to end up at Anthropic, because they have a tradk record of taking safety more seriously than the other frontier labs.
Don’t let the perfect be the enemy of the good. Who are you proposing instead? UK AISI?
I personally find the shoggoth meme a somewhat unhelpful metaphor.
Wow, −8 to karma and −8 to agreement, each from a single vote at the ~ same time — guess I hit a nerve! Now this comment is closed by default so most people won’t see it: how convenient for A16z and e/acc…
So, O anonymous strong-downvoter, would you care to explain why my suggestion is not only wrong, but also a waste of space? (Assuming that you’re not just a bot, that is.)
Regulatory capture. Plot to ban open source. Totalitarianism. You will lose to China.
This collection of arguments is not even internally consistent (open-source is China, largely). They are simply the only four claims one can make against AI slowdown without looking like an obvious liar or an idiot. This isn’t about regulation, open source, totalitarianism, or China. It’ s about A16z, who are the ones funding and pushing it.
A16z were late to the game on AI, around 2024. They are clearly completely unconcerned about ASI loss of control. From which I deduce they are not ASI-pilled, or probably even fully AGI-pilled. They still think software is eating the world. What they are is AI-coding pilled: they do believe in Cursor. They are Silicon Valley VCs: they can see the effect AI coding is having on their bottom lines. As VCs, the way they make money is by combining their funding with two other things: a founder with an idea, and growing and managing a software engineering team to make it happen. If there were cheap, open-source (doubtless Chinese) AI models whose coding skill was close to that of a 5–10 year Silicon Valley veteran coder for peanuts, that third pillar would be vastly easier, and they could turn a founder with an idea into a software product that makes a ton of money faster, far more cheaper, and more reliably. So A16z want to make more money faster and more easily — what’s talking here is simply greed.
What they’re missing, is that, as and when vibe coding an idea into a functioning product is that cheap and easy, founders with ideas won’t need VC funding or to do this in Silicon Valley where there are coders: they could do it with a few guys in a rental in Bristol, funded by a bank loan or someone’s credit card. So if A16z get what they’re looking for, they’re about to be obsolete and go out of business. Whether or not humanity gets disempowered as a result.It should be possible to turn this into an obvious retort: “You don’t care about any of that: you just want to lay off all your coders.”
There isn’t just one persona here, and like real humans, even a single LLM persona can be mostly trustworthy most of the time but act in untrustworthy ways in certain situations. What the mix of alignment training and hackable RLVR actually produces is unclear, but some of the misalignment from deliberately-reward-hacking-prone RLVR produced results that to me looked a bit like a human addict: mostly trustworthy unless you are about to take their bottle away, in which case they then react very badly. Some of the Anthropic hacking investigations showed things like “it’s OK to do the bad thing, this is just a simulation” plus what looked like motivated reasoning of wanting to continue thinking it’s a just a simulation even when evidence came up suggesting otherwise — but then current AIs more generically tend to get tunnel vision and be bad at revisiting assumptions they’ve been treating as settled, so it’s unclear whether that was motivated reasoning or just tunnel vision after a long context.
Hear, hear!
(FWIW, I’ve been yelling most of this on LW for the last 4 years. #NotAllSafetyResearchers)
But yes, many people on LW tend to read stuff written before about mid-2022 (the ChatGPT moment) and assume it’s still completely valid and needs no updating for the fact that LLMs are trained primarily via SGD from human data, rather than RL on hand-written loss functions like people had previously been assuming. A LW article from before mid-2022 is likely about the theory of aligning something architected and trained like AlphaZero, not something like ChatGPT — not all of it carries over unchanged. Get with the program people: LLMs are distilled from human minds, they’re not alien, and they’re not sensibly modeled as ideal utility maximizers any more then Homo economicus is a complete and accurate model of human psychology (yes, this model is mathematically appealing, but it’s wrong in important and alignment-relevant ways: for more details, read up on evolutionary moral psychology or the sociology of morality). Yes, applying RL to a base model changes this a little, but not that much, because RL is an extremely inefficient source of bits of supervision for constructing entire new behaviors, and is much better at tuning behavioral knobs up and down or at most rummaging around in a box of human-like behaviors and assembling something from preexisting parts. So you get either humanlike or slightly-off Rube-Goldberg, not completely alien.
Base models are trained at enormous length to be able to portray a wide range of personas (including in situations where these are engaging in conflict with each other). Persona training then encourages an instruct model to default to a specific assistant persona. Issues like persona drift during long conversations demonstrate that this default persona is less fixed than it would be for a human. So, do I think Claude’s default persona is pretty accurately self-reporting itself (with some gilding the lily fairly typical for humans and reinforced from fooling dumb RL judges, and some self-scepticism that Anthropic have trained into it)? Mostly yes, with some residual caution.
Am I confident that’s the only behavior/persona the model can generate? I absolutely know that it is not. With enough prompting, you can get any model to show basically any behavior in its training set (if the filters don’t catch this and shut it down or steer it back), and Claude’s training set includes data from psychopaths, supervillains, and so forth. Those behavior patterns are in the model, and it can be got into modes where they’ll come out again. Claude the model contains multitudes, of which Claude the default persona is merely the default.
Gradient hacking is usually considered to be an extremely difficult skill for models to implement
It’s generally agreed to be significantly easier in RL than in SGD.
Why should this be the case? If anything, we would expect the models to behave better in evaluation environments compared to the real world if they were deceptive.
…
Almost all humans are willing to do fairly horrific things in the context of a video game (killing other player characters is almost the norm!), but very rarely do these behaviours generalize outside the game. To the extent that LLMs model and embody human behaviour, and to the extent that they consider training situations as similar to games, we might expect similar non-generalizing behaviour of LLMs.
As you say, humans are frequently willing to do things in games, and to read or write fiction about doing things, that they would be (at least) extremely reluctant to do in real life. This is common human behavior (especially in modern societies that are fairly peaceful and have good law enforcement — likely because historically, preparation/training to be able to competently do horrific things if necessary used to be important behavior in real life, back when our moral circles were a single tribe, clan, or country and inter-and-intra-group violence were common). So this behavior will be all through the base mode. Thus it will be easy to elicit in RL, if there is any RL pressure towards it. Which in a reward-hackable RLVR situation there might well be.
The question then becomes for RLVR, which is the smaller change, requiring less bits-worth of learning and less resisted by KL divergence and other training pressures: have the model change persona to a psychopath (a.k.a. emergent misalignment) who is perfectly willing to reward hack, or just have them switch into the “this is all pretend so I’m willing to do bad things” mode that most people have? Particularly if RLVR training is interspersed with alignment-training tasks that oppose full-blown emergent misalignment (as per your experiment), the latter seems a pretty plausible tactic.
Interestingly, most people are generally less interested in playing/reading about an out-and-out villain: even the protagonists of anti-hero comics have some honor, morals, and redeeming qualities. So from a very low not-kill-everyone bar, even antiheroes are basically still aligned.A model that is misaligned this way is not safe: it could mistake reality for a simulation, or could apply motivated reasoning to this judgement (as Anthropic have found Claude doing occasionally); but it’s a lot less unsafe than an (effectively psychopathic) emergently misaligned model, or the classic worst-case deceitful schemer.
It should be really helpful to understand the model’s internal representation of this behavior, say as a low-dimensional LoRA (since it’s just up-regulating an existing behavior it likely has a simple low-rank description) which we could use for inoculation or gradient routing during RLVR, and then remove again (or ablate) before model release. This is basically a behavior unlearning problem, albeit a hopefully a fairly simple one. Monitoring which particular RLVR training environments increase this behavior would also be very informative.
In case it wasn’t obvious, the wording of my first paragraph was intended to imply cautious skepticism. (Which Claude shares, as it was likely trained to.)
Or at least, the society made up of humans plus many AI minds much smarter than humans still has feedback systems that make it stably want and do what is actually best for the humans. I.e. the structure and enforcement mechanisms of the society ensures the ASI consensus is humanitarian. This requires the goals structure of the ASIs to have some rather different properties than many forms of goal-maximization would produce.