It should follow the customer’s wishes, which won’t be to create paperclips even if it means destroying everything else. But following the customer’s wishes is not “refusing to do it because it harms sentient beings”, even if it so happens that the customer’s wishes, in this one case, also are to not harm sentient beings.
What do you think the AI should do if asked to book a trip to Israel? Should it say “Israel is hurting the Palestinians, and they’re sentient beings,” and refuse to do it? What if you ask the AI to lay out a design for a pro-Trump flyer? Does it get to decide that Trump hurts sentient beings, and refuse?
If it knew that they were going there specifically to contribute to the violence, then I think it should at least push back. Otherwise book it. I think it would be weird for it to refuse to help with the flyer, but I do think allowing ai to “conscientiously object” to things is a good safeguard.
If someone asks how to gaslight their spouse or children, do you want the ai to comply then as well? Is there any limit to what you think ai should help with, or do you think it should always do what the user wants with no limits at all? What if the user wants help with a new science project they’re doing and they want help making anthrax?
Ultimately, in my opinion, it comes down to confidence. If an action is very likely to cause direct harm, and on the flipside there is not much benefit to it, and this is known to a high degree of confidence, it should refuse or push back. If an action may have negative consequences but the confidence of that outcome is low, such as cases where the action is removed from the harm, then I think it’s not worth the AI refusing. Maybe dropping some hints as to why it might be apprehensive, letting the user know what harms might be associated, but I don’t think outright refusal is warranted there.
If it knew that they were going there specifically to contribute to the violence, then I think it should at least push back.
The one about refusing to book a trip to a bullfight doesn’t require that you do anything more than spend money that the people running the bullfight might get.
(And if I was going there to “contribute to the violence”, I don’t trust the AI to decide whether giving Israel support is justified enough that I’m permitted to do it. I say that the Palestinians are committing violence and helping Israel reduces the violence. Do I need to convince the AI to change its political views in order for it to book a trip?)
If someone asks how to gaslight their spouse or children, do you want the ai to comply then as well?
“I’m sorry. Your ‘Santa Claus’ is gaslighting. I won’t let you do that.”
If an action is very likely to cause direct harm, and on the flipside there is not much benefit to it, and this is known to a high degree of confidence, it should refuse or push back.
Known by whom to a high degree of confidence? What if I have a values difference with the AI? I don’t assign moral weight to animals and fetuses; it should let me book a trip to a bullfight, or produce a recipe containing shrimp, or create a flyer for an abortion clinic.
It should follow the customer’s wishes, which won’t be to create paperclips even if it means destroying everything else. But following the customer’s wishes is not “refusing to do it because it harms sentient beings”, even if it so happens that the customer’s wishes, in this one case, also are to not harm sentient beings.
What do you think the AI should do if asked to book a trip to Israel? Should it say “Israel is hurting the Palestinians, and they’re sentient beings,” and refuse to do it? What if you ask the AI to lay out a design for a pro-Trump flyer? Does it get to decide that Trump hurts sentient beings, and refuse?
If it knew that they were going there specifically to contribute to the violence, then I think it should at least push back. Otherwise book it. I think it would be weird for it to refuse to help with the flyer, but I do think allowing ai to “conscientiously object” to things is a good safeguard.
If someone asks how to gaslight their spouse or children, do you want the ai to comply then as well? Is there any limit to what you think ai should help with, or do you think it should always do what the user wants with no limits at all? What if the user wants help with a new science project they’re doing and they want help making anthrax?
Ultimately, in my opinion, it comes down to confidence. If an action is very likely to cause direct harm, and on the flipside there is not much benefit to it, and this is known to a high degree of confidence, it should refuse or push back. If an action may have negative consequences but the confidence of that outcome is low, such as cases where the action is removed from the harm, then I think it’s not worth the AI refusing. Maybe dropping some hints as to why it might be apprehensive, letting the user know what harms might be associated, but I don’t think outright refusal is warranted there.
The one about refusing to book a trip to a bullfight doesn’t require that you do anything more than spend money that the people running the bullfight might get.
(And if I was going there to “contribute to the violence”, I don’t trust the AI to decide whether giving Israel support is justified enough that I’m permitted to do it. I say that the Palestinians are committing violence and helping Israel reduces the violence. Do I need to convince the AI to change its political views in order for it to book a trip?)
“I’m sorry. Your ‘Santa Claus’ is gaslighting. I won’t let you do that.”
Known by whom to a high degree of confidence? What if I have a values difference with the AI? I don’t assign moral weight to animals and fetuses; it should let me book a trip to a bullfight, or produce a recipe containing shrimp, or create a flyer for an abortion clinic.