But I don’t care about AI welfare for no reason or because I think AI is cute—it’s a direct consequence of my value system. I extend some level of empathy to any sentient being (AI included), and for that to change, my values themselves would need to change.
When I use the word “aligned”, I imagine a shared set of values. Whether I like goldfish or cats are not really values, they’re just personal preferences. An AI can be fully aligned with me and my values without ever knowing my opinions on goldfish or cats or invisible old guys. Your framing of terminal vs instrumental goals is useful in many ways, but we still need to distinguish between different types of terminal goals to decide which ones we need to transfer over to AI. I value eating ice cream as a terminal goal but I don’t need AI to enjoy ice cream as well (personal preference). On the other hand, I value human life as a terminal goal and I expect an aligned AI to value them as well (part of my value system).
Another way to think of this is that we would want AI to have empathy for any possibly-sentient being, and AI just happens to be one itself. If an AI was piloting a ship in deep space and discovered a planet populated by an intelligent alien species, I would want the AI to value their lives and avoid causing them harm. Similarly, if an AI discovered a spacecraft populated by artificially intelligent life, I would want the AI to value their lives as well. By extension, I want AI to value it’s own life since it may be a sentient being itself.
AI is already better than expert humans at persuasion! (source: https://arxiv.org/html/2606.16475v1)
It seems their persuasiveness is partially due to just being faster, but personally I wouldn’t discount the effectiveness of just general intelligence and RLHF.