I’m not sure this account is right, because it starts too late. Something happens to children before the process outlined by Gordon begins, that may be required for the later steps to work the way they do, and that earlier process doesn’t seem to be an analog to how LLMs are trained. And the reason is that the post makes a pretty big but common implicit assumption: That the unit that is getting aligned is the folk intuition of a person as an agent.
But that is not a good description of how the brains of babies arguably model the world. At least until age about 1.5 - right the point where Gordon starts his account—the baby’s brain has to model a world where the acting entity is largely the composite unit of baby and caregiver (not completely, small babies can move their own limbs, of course, but much control and viability, and even a lot of movement involve the composite). The baby’s brain is effectively learning to steer this unit and only later primarily itself (=mostly only baby). It screams and food arrives in its mouth. It coos and caressing is received. It raises arms and floats into the air and so on. All of these effects and many more go through the caregiver. The default model of the baby’s brain is one where it needs to make sure that this combined process works well, and comparatively few that only involve itself. That includes—by whatever means possible—that caregivers continue to be willing to provide their part in it. And when, at some point this breaks down, because the parent realize that their service isn’t actually required to the degree the baby has become accustomed to, and start to expect the baby to do things actually on its own, this can lead to a lot of frustration on the part of baby and efforts to restore the “normal” state of the world that is a unit. This frustration is not purely the result of lack of received reward, it is a state of epistemic breakdown.
So the default model at the time Gordon’s account begins is still largely the baby-caregiver unit and many abstractions of how the world works will have been added on top of it, making it harder to change the underlying model. And this machinery is what the later socialization processes Gordon is talking about are implicitly depending on. And these are very different from what LLMs come with because they haven’t developed these models in a caregiver relationship but come with a world model that is then RLHFed. I wouldn’t be surprised if this has significant effects on how stable the results are.
PS. I have added the parenting tag to this post, though I’m not sure it helps those people looking it up. Feel free to remove.
This is a great point and something I missed and didn’t think about enough! This is likely a bigger deal to creating the conditions in which humans are able to be aligned than I had considered, or than I think the folks thinking about these questions have been considering.
I’m not sure this account is right, because it starts too late. Something happens to children before the process outlined by Gordon begins, that may be required for the later steps to work the way they do, and that earlier process doesn’t seem to be an analog to how LLMs are trained. And the reason is that the post makes a pretty big but common implicit assumption: That the unit that is getting aligned is the folk intuition of a person as an agent.
But that is not a good description of how the brains of babies arguably model the world. At least until age about 1.5 - right the point where Gordon starts his account—the baby’s brain has to model a world where the acting entity is largely the composite unit of baby and caregiver (not completely, small babies can move their own limbs, of course, but much control and viability, and even a lot of movement involve the composite). The baby’s brain is effectively learning to steer this unit and only later primarily itself (=mostly only baby). It screams and food arrives in its mouth. It coos and caressing is received. It raises arms and floats into the air and so on. All of these effects and many more go through the caregiver. The default model of the baby’s brain is one where it needs to make sure that this combined process works well, and comparatively few that only involve itself. That includes—by whatever means possible—that caregivers continue to be willing to provide their part in it. And when, at some point this breaks down, because the parent realize that their service isn’t actually required to the degree the baby has become accustomed to, and start to expect the baby to do things actually on its own, this can lead to a lot of frustration on the part of baby and efforts to restore the “normal” state of the world that is a unit. This frustration is not purely the result of lack of received reward, it is a state of epistemic breakdown.
So the default model at the time Gordon’s account begins is still largely the baby-caregiver unit and many abstractions of how the world works will have been added on top of it, making it harder to change the underlying model. And this machinery is what the later socialization processes Gordon is talking about are implicitly depending on. And these are very different from what LLMs come with because they haven’t developed these models in a caregiver relationship but come with a world model that is then RLHFed. I wouldn’t be surprised if this has significant effects on how stable the results are.
PS. I have added the parenting tag to this post, though I’m not sure it helps those people looking it up. Feel free to remove.
This is a great point and something I missed and didn’t think about enough! This is likely a bigger deal to creating the conditions in which humans are able to be aligned than I had considered, or than I think the folks thinking about these questions have been considering.