@Jiro it sounds like you don’t believe in transformative AI coming soon? I’m not worried about AIs acting on behalf of humans I’m worried about aligning the AIs values themselves. Our biggest concern with all this is the AI itself decides to kill all sentient beings (including humans). We think the way it acts towards animals now is a good test of how it will act towards humans later. Hence, this is a metric we should be measuring now so we can at least argue how best to address it rather then pretending the metric doesn’t exist.
Even if you think AI will be intelligent, if “you” attempt to align them, it won’t be you. It will be the groups I allude to above. It’s their ideas of harm that get programmed into the AI. If the AI refuses to book a bullfight, it’s not aligned to me; it’s aligned to some animal rights activist.
Tell me, what should an AI do if asked to book a trip to Israel, given that some people think that Israel is causing unjustified harm to sentient beings? Should it refuse on those grounds?
@Jiro it sounds like you don’t believe in transformative AI coming soon? I’m not worried about AIs acting on behalf of humans I’m worried about aligning the AIs values themselves. Our biggest concern with all this is the AI itself decides to kill all sentient beings (including humans). We think the way it acts towards animals now is a good test of how it will act towards humans later. Hence, this is a metric we should be measuring now so we can at least argue how best to address it rather then pretending the metric doesn’t exist.
Even if you think AI will be intelligent, if “you” attempt to align them, it won’t be you. It will be the groups I allude to above. It’s their ideas of harm that get programmed into the AI. If the AI refuses to book a bullfight, it’s not aligned to me; it’s aligned to some animal rights activist.
Tell me, what should an AI do if asked to book a trip to Israel, given that some people think that Israel is causing unjustified harm to sentient beings? Should it refuse on those grounds?