I don’t think that theory [social status] explains quite as much as those people think it does
Do you have a link/explanation for this? I think this may be fairly cruxy, because I’m guessing your intuitions for truth-seeking disagreeable nerd AGI are substantially based on truth-seeking disagreeable nerd humans, so it matters what those humans’ real motivations are.
One line of thought I have here is, there are lots of things such a human or AGI could disagree or talk about or have an interest in, how does it pick which one? I think for the human it probably comes down to some kind of subconscious status calculation, but in either case, how does the AGI do it if it doesn’t have its own status motivations or other long-term goals?
One line of thought I have here is, there are lots of things such a human or AGI could disagree or talk about or have an interest in, how does it pick which one? I think for the human it probably comes down to some kind of subconscious status calculation, but in either case, how does the AGI do it if it doesn’t have its own status motivations or other long-term goals?
I think we do want the AGI to have long-term goals, just not to have exclusively long-term goals, at least in the §6.2.1 approach. See my old 2021 post Consequentialism & corrigibility. Again, if you or me is the model to be inspired by, then I assume we both care about the future of life being great (long-term goal), but we both also enjoy figuring things out in the here and now, and we both also have principles that we take pride in. And for my part, I wouldn’t want to be benevolent dictator of the universe even if I could, that’s way too much responsibility, sounds terrifying.
To be clear, I’m generally expecting the process to be kinda messy, where it’s kinda hard to reason about where the AGI winds up. From my perspective, Step 1 is to have any plan at all that could plausibly work, and then we can move on to making it easier to test and de-risk the plan in advance, to the extent possible. As mentioned at the bottom, I’m still hard at work trying to get more clarity, to the extent possible, in order to make the process of AGI motivation development more predictable and legible and less messy.
I don’t think that theory [social status] explains quite as much as those people think it does
Do you have a link/explanation for this? I think this may be fairly cruxy, because I’m guessing your intuitions for truth-seeking disagreeable nerd AGI are substantially based on truth-seeking disagreeable nerd humans, so it matters what those humans’ real motivations are.
Not all in one place, but here’s some pointers (and feel free to ask follow-ups).
Since the works of Robin Hanson are popular on this forum, I will say a bit more about where I differ from Elephant In The Brain. My biggest complaint is the part where they say:
As we mentioned earlier, people are profoundly ignorant about laughter’s meaning and purpose (at least in our default state, before learning the science). But where does this ignorance come from? Why does introspection fail us so spectacularly here?
It’s not simply because laughter is involuntary, outside our conscious control. Flinching, for example, is also involuntary, and yet we understand perfectly well why we do it: to protect ourselves from getting hit. Thus our ignorance about laughter needs further explanation.
I disagree that it “needs further explanation”. I think we start out ignorant of literally everything, until we learn it / figure it out. And I think that figuring out the evolutionary purpose of laughter is just inherently much harder than figuring out the evolutionary purpose of flinching. It’s less obvious / salient, for various reasons that I claim are pretty obvious if you think about it. I don’t think there’s any more to it than that.
I also don’t think there can be more to it than that. To explain what I mean by that, imagine if I said: “Here’s the source code for training an image-classifier ConvNet from random initialization using uncontrolled external training data. Can you please edit this source code so that the trained model winds up confused about the shape of Toyota Camry tires specifically?” The answer is: “Nope. Sorry. There is no possible edit I can make to this PyTorch source code such that that will happen.” By the same token, even if, as that book argues, there is a strong evolutionary pressure to make humans specifically confused about the evolutionary purpose of laughter, I don’t think there is any possible genetic change that would make that happen. Related discussion here.
Up a level, this strategic self-deception idea comes out of the “evolved modularity” framework in evolutionary psychology, and I reject that whole broader framework as well. See §1.1 of “My take on Jacob Cannell’s take on AGI safety” for the background, and “Learning from scratch” in the brain for why I don’t buy it (kinda related to the thing above about PyTorch).
For related reasons, I reject the idea that “status” (per se) could possibly be an innate goal. It’s just too abstract. Copying from here:
Explaining how human social instincts work is tricky mainly because of the “symbol grounding problem”. In brief, everything we know—all the interlinked concepts that constitute our understanding of the world and ourselves—is created “from scratch” in the cortex by a learning algorithm, and thus winds up in the form of a zillion unlabeled data entries like “pattern 387294 implies pattern 579823 with confidence 0.184”, or whatever. Yet certain activation states of these unlabeled entries—e.g., the activation state that encodes the fact that Jun just told me that Xiu thinks I’m cute—need to somehow trigger social instincts in the Steering Subsystem. So there must be some way that the brain can “ground” these unlabeled learned concepts.
Now, it’s not that there’s no way to solve the symbol grounding problem to make someone want social status—indeed, that obviously happens, and in my post Neuroscience of human social instincts: a sketch, I attempt to explain how. It’s that there’s no way to solve the symbol grounding problem to make someone want social status specifically. Realistically, the genetic mechanism is just not gonna be that specific. (More on which shortly.)
So here’s where we’re at so far: (1) without Trivers-style self-deception, we have a harder time explaining away interoceptive reports like “That’s not status-seeking because I’m not trying to impress anyone” (we can still try to explain it away, but it’s harder); and (2) the idea that “status” could have a special place as an innate end-goal is implausible anyway.
That post’s §2 is the part that’s closest to status-seeking: people are motivated to have actual interactions where an actual person has positive associations with you. But even in this case, status-seeking is just one of several consequences of the same innate drive. The others are credit-seeking / blame-avoidance, and norm-following / norm-enforcement. I understand that you can try to unify these by hypothesizing that (say) credit-seeking is a means-to-an-end for achieving social status, but that’s just not how it works in my neuroscience model: credit-seeking / blame avoidance, status-seeking, and norm-enforcement all emerge in the same way, at the same level, via the same mechanism.
And then that post’s §3 gets even more distant from the conventional notion of status-seeking, by analyzing how we can feel pride in ourselves and our actions, and how this is another direct consequence of the same innate drive, and not a secret means-to-an-end to achieving social status.
And indeed, it’s easy to come up with cases where pride vs future-status come apart, and where people follow the former over the latter. E.g. people sometimes stand up for principles even if they think everyone will scorn them for it, because they’re following their own moral compass. Certainly the moral compass has something to do with what other people have said and thought over the course of the person’s life, but the relation can be quite indirect (e.g. people may care about how a cartoon character would judge their behavior), and not well-described as “trying to wind up with high social status”, consciously or unconsciously.
Do you have a link/explanation for this? I think this may be fairly cruxy, because I’m guessing your intuitions for truth-seeking disagreeable nerd AGI are substantially based on truth-seeking disagreeable nerd humans, so it matters what those humans’ real motivations are.
One line of thought I have here is, there are lots of things such a human or AGI could disagree or talk about or have an interest in, how does it pick which one? I think for the human it probably comes down to some kind of subconscious status calculation, but in either case, how does the AGI do it if it doesn’t have its own status motivations or other long-term goals?
I think we do want the AGI to have long-term goals, just not to have exclusively long-term goals, at least in the §6.2.1 approach. See my old 2021 post Consequentialism & corrigibility. Again, if you or me is the model to be inspired by, then I assume we both care about the future of life being great (long-term goal), but we both also enjoy figuring things out in the here and now, and we both also have principles that we take pride in. And for my part, I wouldn’t want to be benevolent dictator of the universe even if I could, that’s way too much responsibility, sounds terrifying.
To be clear, I’m generally expecting the process to be kinda messy, where it’s kinda hard to reason about where the AGI winds up. From my perspective, Step 1 is to have any plan at all that could plausibly work, and then we can move on to making it easier to test and de-risk the plan in advance, to the extent possible. As mentioned at the bottom, I’m still hard at work trying to get more clarity, to the extent possible, in order to make the process of AGI motivation development more predictable and legible and less messy.
Not all in one place, but here’s some pointers (and feel free to ask follow-ups).
Let’s start with the Hansonian notion of strategic self-deception (as distinct from plain old motivated reasoning which is real and important), which I assume he got from Robert Trivers. I broadly reject that notion. E.g. here’s a footnote in my post about laughter:
Up a level, this strategic self-deception idea comes out of the “evolved modularity” framework in evolutionary psychology, and I reject that whole broader framework as well. See §1.1 of “My take on Jacob Cannell’s take on AGI safety” for the background, and “Learning from scratch” in the brain for why I don’t buy it (kinda related to the thing above about PyTorch).
For related reasons, I reject the idea that “status” (per se) could possibly be an innate goal. It’s just too abstract. Copying from here:
Now, it’s not that there’s no way to solve the symbol grounding problem to make someone want social status—indeed, that obviously happens, and in my post Neuroscience of human social instincts: a sketch, I attempt to explain how. It’s that there’s no way to solve the symbol grounding problem to make someone want social status specifically. Realistically, the genetic mechanism is just not gonna be that specific. (More on which shortly.)
So here’s where we’re at so far: (1) without Trivers-style self-deception, we have a harder time explaining away interoceptive reports like “That’s not status-seeking because I’m not trying to impress anyone” (we can still try to explain it away, but it’s harder); and (2) the idea that “status” could have a special place as an innate end-goal is implausible anyway.
That brings us to my Social drives 2: “Approval Reward”, from norm-enforcement to status-seeking, which is a follow-up to Neuroscience of human social instincts: a sketch where I attempt to connect the neuroscience to everyday life:
That post’s §2 is the part that’s closest to status-seeking: people are motivated to have actual interactions where an actual person has positive associations with you. But even in this case, status-seeking is just one of several consequences of the same innate drive. The others are credit-seeking / blame-avoidance, and norm-following / norm-enforcement. I understand that you can try to unify these by hypothesizing that (say) credit-seeking is a means-to-an-end for achieving social status, but that’s just not how it works in my neuroscience model: credit-seeking / blame avoidance, status-seeking, and norm-enforcement all emerge in the same way, at the same level, via the same mechanism.
And then that post’s §3 gets even more distant from the conventional notion of status-seeking, by analyzing how we can feel pride in ourselves and our actions, and how this is another direct consequence of the same innate drive, and not a secret means-to-an-end to achieving social status.
And indeed, it’s easy to come up with cases where pride vs future-status come apart, and where people follow the former over the latter. E.g. people sometimes stand up for principles even if they think everyone will scorn them for it, because they’re following their own moral compass. Certainly the moral compass has something to do with what other people have said and thought over the course of the person’s life, but the relation can be quite indirect (e.g. people may care about how a cartoon character would judge their behavior), and not well-described as “trying to wind up with high social status”, consciously or unconsciously.