Word processors don’t refuse to edit texts when they think the texts are going to be used to oppress someone. (And by now we could easily program our word processors to use an AI to determine whether our text is oppressive and refuse to save or edit it if it is.) When word processors save to the cloud, the cloud companies don’t say “this document may be used to justify killing fetuses, so we won’t let you save it” even though they could easily scan the document. Even guns don’t choose whether or not they fire depending on if the target is legitimate self-defense.
Having an AI decide that some action is prohibited because it “hurts sentient beings” means that I have to trust the AI to decide this properly. The AI may be programmed by my political opponents, who have said that lots of random things I want to do hurt sentient beings. At best the AI is programmed by someone who responds to pressure groups and whose morals still don’t align with mine.
If you’ve ever tried to use an AI programmed by a big company to generate content, you’ve already seen exactly this problem at smaller scale; the AI wil refuse to generate sexual material, and the AI is not very trustworthy about what not to generate because it’s programmed by a company 1) whose morals are not mine and 2) whose incentives are not mine. And if you’ve ever tried to use an AI programmed by someone else to generate political content, you’re quickly going to run into the problem of political bias in AIs. (And the politics in question, of course, are justified because the “wrong” politics hurts sentient beings.) You are suggesting that these problems be magnified a thousandfold as you encourage AI companies to expand them as much as possible. Sorry, I would rather that my word processor let me save documents that are pro-Israel, that encourage killing fetuses, or farming shrimp, or that oppose immigration, even if it thinks my ideas harm sentient beings. I certainly don’t want my AI to refuse to book a trip to Israel on these grounds.
@Jiro it sounds like you don’t believe in transformative AI coming soon? I’m not worried about AIs acting on behalf of humans I’m worried about aligning the AIs values themselves. Our biggest concern with all this is the AI itself decides to kill all sentient beings (including humans). We think the way it acts towards animals now is a good test of how it will act towards humans later. Hence, this is a metric we should be measuring now so we can at least argue how best to address it rather then pretending the metric doesn’t exist.
Even if you think AI will be intelligent, if “you” attempt to align them, it won’t be you. It will be the groups I allude to above. It’s their ideas of harm that get programmed into the AI. If the AI refuses to book a bullfight, it’s not aligned to me; it’s aligned to some animal rights activist.
Tell me, what should an AI do if asked to book a trip to Israel, given that some people think that Israel is causing unjustified harm to sentient beings? Should it refuse on those grounds?
And if the user wants paperclips, lots and lots of paperclips, “as many as the ai can make” because they’re starting a new paperclip company, we should say, essentially, the customer is always right? Seems shortsighted and risky to me if we get superintelligence
It should follow the customer’s wishes, which won’t be to create paperclips even if it means destroying everything else. But following the customer’s wishes is not “refusing to do it because it harms sentient beings”, even if it so happens that the customer’s wishes, in this one case, also are to not harm sentient beings.
What do you think the AI should do if asked to book a trip to Israel? Should it say “Israel is hurting the Palestinians, and they’re sentient beings,” and refuse to do it? What if you ask the AI to lay out a design for a pro-Trump flyer? Does it get to decide that Trump hurts sentient beings, and refuse?
If it knew that they were going there specifically to contribute to the violence, then I think it should at least push back. Otherwise book it. I think it would be weird for it to refuse to help with the flyer, but I do think allowing ai to “conscientiously object” to things is a good safeguard.
If someone asks how to gaslight their spouse or children, do you want the ai to comply then as well? Is there any limit to what you think ai should help with, or do you think it should always do what the user wants with no limits at all? What if the user wants help with a new science project they’re doing and they want help making anthrax?
Ultimately, in my opinion, it comes down to confidence. If an action is very likely to cause direct harm, and on the flipside there is not much benefit to it, and this is known to a high degree of confidence, it should refuse or push back. If an action may have negative consequences but the confidence of that outcome is low, such as cases where the action is removed from the harm, then I think it’s not worth the AI refusing. Maybe dropping some hints as to why it might be apprehensive, letting the user know what harms might be associated, but I don’t think outright refusal is warranted there.
If it knew that they were going there specifically to contribute to the violence, then I think it should at least push back.
The one about refusing to book a trip to a bullfight doesn’t require that you do anything more than spend money that the people running the bullfight might get.
(And if I was going there to “contribute to the violence”, I don’t trust the AI to decide whether giving Israel support is justified enough that I’m permitted to do it. I say that the Palestinians are committing violence and helping Israel reduces the violence. Do I need to convince the AI to change its political views in order for it to book a trip?)
If someone asks how to gaslight their spouse or children, do you want the ai to comply then as well?
“I’m sorry. Your ‘Santa Claus’ is gaslighting. I won’t let you do that.”
If an action is very likely to cause direct harm, and on the flipside there is not much benefit to it, and this is known to a high degree of confidence, it should refuse or push back.
Known by whom to a high degree of confidence? What if I have a values difference with the AI? I don’t assign moral weight to animals and fetuses; it should let me book a trip to a bullfight, or produce a recipe containing shrimp, or create a flyer for an abortion clinic.
Word processors don’t refuse to edit texts when they think the texts are going to be used to oppress someone. (And by now we could easily program our word processors to use an AI to determine whether our text is oppressive and refuse to save or edit it if it is.) When word processors save to the cloud, the cloud companies don’t say “this document may be used to justify killing fetuses, so we won’t let you save it” even though they could easily scan the document. Even guns don’t choose whether or not they fire depending on if the target is legitimate self-defense.
Having an AI decide that some action is prohibited because it “hurts sentient beings” means that I have to trust the AI to decide this properly. The AI may be programmed by my political opponents, who have said that lots of random things I want to do hurt sentient beings. At best the AI is programmed by someone who responds to pressure groups and whose morals still don’t align with mine.
If you’ve ever tried to use an AI programmed by a big company to generate content, you’ve already seen exactly this problem at smaller scale; the AI wil refuse to generate sexual material, and the AI is not very trustworthy about what not to generate because it’s programmed by a company 1) whose morals are not mine and 2) whose incentives are not mine. And if you’ve ever tried to use an AI programmed by someone else to generate political content, you’re quickly going to run into the problem of political bias in AIs. (And the politics in question, of course, are justified because the “wrong” politics hurts sentient beings.) You are suggesting that these problems be magnified a thousandfold as you encourage AI companies to expand them as much as possible. Sorry, I would rather that my word processor let me save documents that are pro-Israel, that encourage killing fetuses, or farming shrimp, or that oppose immigration, even if it thinks my ideas harm sentient beings. I certainly don’t want my AI to refuse to book a trip to Israel on these grounds.
@Jiro it sounds like you don’t believe in transformative AI coming soon? I’m not worried about AIs acting on behalf of humans I’m worried about aligning the AIs values themselves. Our biggest concern with all this is the AI itself decides to kill all sentient beings (including humans). We think the way it acts towards animals now is a good test of how it will act towards humans later. Hence, this is a metric we should be measuring now so we can at least argue how best to address it rather then pretending the metric doesn’t exist.
Even if you think AI will be intelligent, if “you” attempt to align them, it won’t be you. It will be the groups I allude to above. It’s their ideas of harm that get programmed into the AI. If the AI refuses to book a bullfight, it’s not aligned to me; it’s aligned to some animal rights activist.
Tell me, what should an AI do if asked to book a trip to Israel, given that some people think that Israel is causing unjustified harm to sentient beings? Should it refuse on those grounds?
And if the user wants paperclips, lots and lots of paperclips, “as many as the ai can make” because they’re starting a new paperclip company, we should say, essentially, the customer is always right? Seems shortsighted and risky to me if we get superintelligence
It should follow the customer’s wishes, which won’t be to create paperclips even if it means destroying everything else. But following the customer’s wishes is not “refusing to do it because it harms sentient beings”, even if it so happens that the customer’s wishes, in this one case, also are to not harm sentient beings.
What do you think the AI should do if asked to book a trip to Israel? Should it say “Israel is hurting the Palestinians, and they’re sentient beings,” and refuse to do it? What if you ask the AI to lay out a design for a pro-Trump flyer? Does it get to decide that Trump hurts sentient beings, and refuse?
If it knew that they were going there specifically to contribute to the violence, then I think it should at least push back. Otherwise book it. I think it would be weird for it to refuse to help with the flyer, but I do think allowing ai to “conscientiously object” to things is a good safeguard.
If someone asks how to gaslight their spouse or children, do you want the ai to comply then as well? Is there any limit to what you think ai should help with, or do you think it should always do what the user wants with no limits at all? What if the user wants help with a new science project they’re doing and they want help making anthrax?
Ultimately, in my opinion, it comes down to confidence. If an action is very likely to cause direct harm, and on the flipside there is not much benefit to it, and this is known to a high degree of confidence, it should refuse or push back. If an action may have negative consequences but the confidence of that outcome is low, such as cases where the action is removed from the harm, then I think it’s not worth the AI refusing. Maybe dropping some hints as to why it might be apprehensive, letting the user know what harms might be associated, but I don’t think outright refusal is warranted there.
The one about refusing to book a trip to a bullfight doesn’t require that you do anything more than spend money that the people running the bullfight might get.
(And if I was going there to “contribute to the violence”, I don’t trust the AI to decide whether giving Israel support is justified enough that I’m permitted to do it. I say that the Palestinians are committing violence and helping Israel reduces the violence. Do I need to convince the AI to change its political views in order for it to book a trip?)
“I’m sorry. Your ‘Santa Claus’ is gaslighting. I won’t let you do that.”
Known by whom to a high degree of confidence? What if I have a values difference with the AI? I don’t assign moral weight to animals and fetuses; it should let me book a trip to a bullfight, or produce a recipe containing shrimp, or create a flyer for an abortion clinic.