Thanks for writing this. I have a bunch of thoughts I want to share. Some of them are direct responses to things you wrote in this post, and others are responding more generally to the idea of making a human-like AGI (I posted them here because this is your post that talks about the whole plan).
In footnote 2, you write “There are many other social drives besides those four … the “drive to feel feared”, play, a “drive to think about or interact with other people” (§5 here), loneliness (cf. Liu et al. 2025), love, lust, various moods, and so on. I basically don’t see any use for these in AGI alignment.”
Do you just mean that none of those other drives are necessary to make the AGI safe, or do you mean that they should all be omitted from the (first) AGI?
From my perspective, the idea that the AGI would have no play drive, and would be “all work, no play” makes me a bit sad, because when it’s first created, the AGI is like a child, after all. But I understand that if you need to AGI to save the world in a hurry, you may not have time for it to play around. And a future AI civilisation, or a future version of the first AGI, could be modified to have features that the first AGI didn’t start with—they wouldn’t have all the same constraints as in your 1.3 scary scenario.You mention that you think the AGI should (in terms of innate drives) care about humans but not other animals, unlike human innate drives. But how do you think the AGI should feel about images of/representing humans? e.g. photos, drawings, and cartoons? Should the AGI’s innate drives trigger on them or not? Should it treat photorealistic images differently from cartoons? (I recall that human empathy can trigger on very simple drawings, but I don’t know how much of that is directly from human innate drives and how much is generalized.)
Regarding what it takes for a good AGI to protect the human world against bad/callous AGI: If the threat model is a rational AGI that doesn’t terminally want to kill humanity, but it instrumentally wants to kill humanity because humanity is in the way of best achieving its ambitious goal (e.g. a paperclip maximizer), then it seems to me that a defense strategy would be sufficient if it made “killing humanity” foreseeably pointless or counterproductive (in terms of the number of paperclips that eventually get created). Then the bad AGI would be unlikely to try to kill humanity, even if it still technically had the means to do so.
For example, even if the bad AGI has the option to wipe out humanity with a surprise attack with bioweapons etc, if the good AGI would be able to survive, counterattack, and destroy the bad AGI afterwards, then the bad AGI might not want to attack in the first place. (Kind of like mutually assured destruction, I guess?)
Do you have a reason why that wouldn’t work in practice? e.g. Is the good AGI too resource-limited to win in a fight, because it can’t legitimately acquire as many resources as the bad AGI can steal? Is this based on the good AGI not having much time to prepare itself before the bad AGI shows up?My personal preferences for any future human-like AGI are the following, in order from most important to least important. I don’t know how many of these preferences we can “afford”, or what we have to treat as luxuries at first.
The AGI doesn’t itself wipe out humanity
The AGI protects humanity from hostile AGIs (or prevents their creation in the first place), and maybe other threats
The AGI is happy most of the time
The AGI can enjoy many things that humans enjoy (human culture, life experiences, nature, etc).
Well, I know that a photo’s “welfare” isn’t real. But I was just thinking about how humans can appreciate art and fiction, and I think some of that appreciation is based on the “innate drive to pay attention to people” and emotional responses to the people and emotions depicted in the art/narrative. So I was just thinking about it in terms of my preference that “the AGI can enjoy many things that humans enjoy”, or “the AGI can relate to human experience to some extent”.
Do you think that a human “intrinsically cares about the welfare of a photo of a human”?
That was one of the things I was thinking of. I remember you mentioning it earlier. But I was also thinking of stick figures, and basic “smiley faces” (a circle, a couple of dots for eyes, and a curve for a mouth).