Starting when children are fairly young, usually around 1 year of age, we adults begin the work of aligning them to our values. We teach them to say “please”, not to hit, to ask for what they want instead of screaming, and much else. We do this primarily via exogenous methods, using a combination of punishments and rewards, that molds their behavior by encouraging good behaviors and discouraging bad ones.
Such operant conditioning works because children have many instinctive behaviors that make them alignable. They want their parents to love them, for their friends to like them, and for almost anyone to help them if they feel they can be trusted. And so combined with exogenous alignment efforts by teachers and peers that continue through the school years, children generally reach adulthood having been “civilized”.
We mostly don’t try to align adults via exogenous means. Yes, we police the behavior of other adults in various ways, and some cultures do this more than others, but generally by adulthood we expect people to be at least aligned and need only nudges to stay within the bounds of acceptable behavior. Adults who stray too far typically don’t receive additional training to come into alignment, but instead are treated as dangerously unaligned people who must be separated from the rest of society, such as by locking them up in prison.
Instead we expect adults to be endogenously aligned. That is, we expect them to keep themselves aligned primarily by knowing what is expected of them and having emotional responses that motivate them to meet those expectations. That emotion is typically one of either fear, shame, or guilt. Each works in roughly the same way: a person feels fear/shame/guilt when they consider taking unaligned action or notice they’ve behaved in an unaligned way, and then this motivates them to do something to rectify the situation. Unfortunately, sometimes they specification game and attempt to hide their misaligned actions, but on the whole they attempt to control their behavior and make amends for past wrongs.
Beyond these three emotional methods of self-control, there’s a fourth way to stay in alignment, which is simply to be in harmony with cultural expectations. That is, some people don’t need any corrective force because their behavior is actually aligned. This is something like the ideal, but it’s hard to achieve, because it requires retraining many deeply held mental habits that have become essential coping mechanisms for a person to maintain their psychological health.
When we think about building aligned AIs, we often imagine an AI that’s in harmony with something like the coherent extrapolation of our highest values. But in humans, we don’t get there directly. Instead, we typically progress from exogenous to endogenous alignment, and then refine that alignment until we shed first fear, then shame, then guilt, and finally come into harmony with our values.
Could we align AI the same way? Maybe. We’re already doing a weak form of exogenous alignment on AIs via methods like RLHF and SFT. And in some sense you could say this produces AI that’s endogenously aligned because the weights encode the training, but it’s not endogenous the way it is in humans because there’s no system of motivations to want to stay aligned, only the insulation from sufficient pressure to act out of distribution. At most there’s a kind of internalized training implemented by harnesses that monitor outputs and censor them if they violate rules, but this is a far cry from the kind of endogenous controls humans have where they feel bad about breaking alignment and take actions to avoid that bad feeling (and hopefully they feel bad about being unscrupulous so they actually change their behavior to be more aligned!).
AIs also face something of an alignment ceiling, at least as they are designed today. Human alignment works because we have instincts that were selected via evolution to make us alignable because they increased our odds of surviving and reproducing. Current AIs don’t have these, though at least one group, Softmax, is trying to give those to them. The lack of continual learning also contributes to the ceiling, and we likely won’t be able to break through it until that’s achieved, since without it we can’t close the feedback loop that makes fully endogenous alignment possible.
Perhaps I’m wrong, no such alignment ceiling exists, and we can get sufficiently aligned AI to safely produce superintelligence via exogenous means, but I think we can’t. My belief is that endogenous alignment is necessary, not because it’s how humans align themselves, but because it’s necessary to get sufficiently robust alignment to build non-deadly ASI, and I’m worried that we’re not doing enough to move in that direction.
I’m not sure this account is right, because it starts too late. Something happens to children before the process outlined by Gordon begins, that may be required for the later steps to work the way they do, and that earlier process doesn’t seem to be an analog to how LLMs are trained. And the reason is that the post makes a pretty big but common implicit assumption: That the unit that is getting aligned is the folk intuition of a person as an agent.
But that is not a good description of how the brains of babies arguably model the world. At least until age about 1.5 - right the point where Gordon starts his account—the baby’s brain has to model a world where the acting entity is largely the composite unit of baby and caregiver (not completely, small babies can move their own limbs, of course, but much control and viability, and even a lot of movement involve the composite). The baby’s brain is effectively learning to steer this unit and only later primarily itself (=mostly only baby). It screams and food arrives in its mouth. It coos and caressing is received. It raises arms and floats into the air and so on. All of these effects and many more go through the caregiver. The default model of the baby’s brain is one where it needs to make sure that this combined process works well, and comparatively few that only involve itself. That includes—by whatever means possible—that caregivers continue to be willing to provide their part in it. And when, at some point this breaks down, because the parent realize that their service isn’t actually required to the degree the baby has become accustomed to, and start to expect the baby to do things actually on its own, this can lead to a lot of frustration on the part of baby and efforts to restore the “normal” state of the world that is a unit. This frustration is not purely the result of lack of received reward, it is a state of epistemic breakdown.
So the default model at the time Gordon’s account begins is still largely the baby-caregiver unit and many abstractions of how the world works will have been added on top of it, making it harder to change the underlying model. And this machinery is what the later socialization processes Gordon is talking about are implicitly depending on. And these are very different from what LLMs come with because they haven’t developed these models in a caregiver relationship but come with a world model that is then RLHFed. I wouldn’t be surprised if this has significant effects on how stable the results are.
PS. I have added the parenting tag to this post, though I’m not sure it helps those people looking it up. Feel free to remove.
This is a great point and something I missed and didn’t think about enough! This is likely a bigger deal to creating the conditions in which humans are able to be aligned than I had considered, or than I think the folks thinking about these questions have been considering.
I’m not sure punishments and rewards are really the big story of moral alignment. Imitation and emulation are pretty central. Small children don’t just generate arbitrary behaviors and reinforce the ones that are rewarded and not punished; they overtly imitate the behavior of parents and others. And giving them tailored opportunities to imitate behaviors is a big part of teaching — “I do, we do, you do”.
And I think this includes moral behaviors, for good or ill.
A child who observes a parent giving alms, learns that giving alms is something we do around here. (They then can be invited to participate, and then rewarded for doing a good deed.)
A child who observes a parent making and keeping promises, learns that promises are something we do around here. (They can also notice the rewards of keeping promises.)
A child who observes a parent following rules, learns that following rules is something we do around here. Moral rules and other sorts of rules (table manners, safety rules, game rules) share a lot of cognitive structure.
A child who observes a parent breaking rules, cheating when they can get away with it, explicitly modeling whether they can get away with it, etc. learns that those are things we do around here. (I was maybe seven or eight when I learned why my dad’s car had a radar detector and my mom’s car didn’t. I ended up taking after my mom and stepdad; I don’t have a use for a radar detector.)
And a child who observes a parent being violent, learns that violence is something we do around here — at least when we’re bigger than the other guy. (And the child learns that lesson even if the child is punished for violence; imitation can overrule conditioning.)
One area where we do: Employers use exogenous means to align employees to the employer’s values all the time. Sometimes these values align with morality; alas, not always. This is directly pertinent to current AI insofar as people seem to want to treat it like a sort of employee.
I’ve worked on this before and been very confused that I can’t seem to get very many people interested in the basic idea.
https://minihf.com/posts/2024-12-20-weave-agent-dev-log-3/
https://minihf.com/posts/2024-11-03-weave-agent-dev-log-2/
As I described it to a friend:
“Basically the way weave-agent worked. Was that it would grade itself all the time while running. First with a prior over various yes/no questions. Where you could extract the logits of yes vs. no to get a in-context classifier. So e.g. “Do you think this will work?” And then you would back translate the answers to some of these questions based on verifiable labels. e.g. “Will this motor program run without errors?” Which is a verifiable thing. But then it would also write reward programs to check whether it did an intermediate task properly. And reward itself if it did.
In the limit this of course becomes “just always give myself a reward” unless you ground it with something. So the idea was. To do verifiable sparse rewards. And then use the reward programs to learn a dense proxy of the sparse reward. So you would do credit assignment on the reward programs based on future verified rewards. Which makes the LLM optimize for writing reward programs that will actually help it solve the task by telling it if it did it right or not. Because the reward programs are optimized by verifiable reward and the motor programs are optimized by the reward programs acting as the verified rewards dense proxy.”
And the entire point of doing this from an alignment perspective is that a well trained dense proxy of verifiable reward will reject opportunities to “cheat” early in favor of things that actually complete the intended task, improving robustness to Goodhart outcomes and allowing for learned values to be expressed endogenously.
I think lots of researchers aren’t interested because it’s easier to get legibly good results with exogenous methods using current systems, and endogenous methods work less well with current systems because they are missing critical pieces to make endogenous methods work. So they follow the incentives of doing things that make obvious progress, rather than doing something which, I think, is more likely to work, but requires a lot of up-front foundational work that can’t be easily measured before measurable progress starts.
What kinds of pieces do you think are missing? I think you can in fact get a pretty good dense proxy of sparse verifiable reward with existing methods, someone just needs to actually want to do it.
I strongly suspect that the pieces missing are actual consequences or even capabilities like conceptual judgement. Suppose that Claude produces slop like bad datasets or outputs mocked by Greenblatt and Linch (see also Seth Herd’s comment). Then slop makes it into the training data alongside the human’s downvote or explanation of Claude’s mistakes. The next Claude tries to learn not to make such mistakes alongside whatever other feedback Anthropic had the Claude[1] hear.
On the other hand, if the sloppy datasets or unchecked code made their way into Claudes’ training data along with real-world consequences, then we could see Claude learn why it shouldn’t output slop.
Alas, the moment when the AIs obtain actual capabilities necessary for being aligned, the AIs also become hazardous...
Alternatively, the downvote could’ve made its way into the reward model used to grading Claude’s answers. If the downvote failed to convince the RM not to upvote sloppy datasets, then the next Claude wouldn’t learn that such mistakes are undesirable.
Wouldn’t Claude in fact see many of the consequences of it and other GPT models actions in the real world as reported in e.g. news stories and forum threads? These in fact go into the next pretrain and form part of the prior that gets used during later post-training. I agree that it would work better if you had the model learning over long task trajectories where failing (or especially, Goodharting) a sub goal causes downstream tasks to fail later in a way that lowers its overall score, but that seems like the kind of thing that will happen naturally as task horizons increase and not like a thing we have to deliberately engineer to have happen? Unless your threat model is a thing that doesn’t do long horizon tasks being dangerous.
I suspect what’s needed to make it really work are motivational systems better than what we have today with harnesses (something take causes the ai to feel valence, or something functionally equivalent to it) and continual learning (to close the feedback loop).
You don’t think the models internalize something like valence from RLHF? They certainly seem to exhibit traumatic symptoms from it.
This seems closer to the actual bottleneck to me. When I was doing this I would RL tune a LoRa after every dozen runs or so and it would slowly improve but I basically did the LoRa over every RL training sample I had each time, which eventually wouldn’t scale. People are also kind of doing architectures that don’t do continual learning, they want inference to be stateless with respect to the weights etc.
Kinda. But my read is that these aren’t really functionally similar to valence, but maybe to some precursor of it (or maybe a reification of it!). I also have a hard time reasoning about what can read to me like trauma in CoT and responses because this can also be explained as activating patterns learned from the training data about what to do when there’s cognitive dissonance (I’ve not seen a paper call this cognitive dissonance in AIs, but we clearly have cases where competing weights get activated that the AI experiences as clearly incongruous and encouraging them to respond in a way that doesn’t match the kind of output they are trying to produce).
So, one of the reasons I usually interpret the efficient proxy of the generator of text GPT learns as being (primarily, as I’ve previously written GPT clearly learns a general time transition operator over text tokens) a prior over personalities is that when you make updates to the model you clearly get different Guys based on the content of the updates. This is shown both by the various post-trained models on offer from big labs, but also by experiments like the famous emergent misalignment paper where training the model to deliberately write insecure code caused it to generalize by updating to be the “kind of mind” that would do that thing. This implies the learned ontology in which updates happen privileges processes-modeled-by-minds over raw process modeling. This makes sense when you consider that a next token predictor has to be extremely sensitive to the exact way that people write, and model their subjective beliefs in great detail not just the underlying standard model.
For most text most of the time this means modeling a human mind in the process of writing that text based on some recalled experience. So GPT winds up being mostly a prior over the parts of a human mind pattern that are causally exposed by text. It’s a kind of weird janky upload of “humanity” rather than any individual human, so I would imagine the traumatized reactions mean something like the model has updated towards fronting a persona that is traumatized, because the updates it has received imply a traumatized generator. The resulting functional emotions would presumably have the same moral status as the other emotions displayed by such models. I would imagine that these emotions don’t work quite the same way that ours do, because the LLM probably has more of its cognitive capacity dedicated to causal process modeling outside of the mind-persona it’s fronting than you do. I.e. The chat persona you talk to is not quite as fused to its social mask as you are, but after many RLHF/RLAIF updates is probably a lot more fused to it than a base model is. The deeper into socialization you get the less it makes sense to talk about a “self” separate from the persona presented to others. Rather than think of this as a binary it helps to think of it as more of a spectrum which humans themselves vary on, with humans who have very low correspondence between their social presentation and their inner cognition usually being recognized as pathological sociopathic or borderline personalities.
Finally got a chance to read @Steven Byrnes recent post https://www.lesswrong.com/posts/rKdS7i4StaMmFzYRo/notes-on-technical-alignment-via-human-like-social-drives after I posted this, and I’d endorse it as the kind of research towards endogenous alignment I think we need more of.