And if the user wants paperclips, lots and lots of paperclips, “as many as the ai can make” because they’re starting a new paperclip company, we should say, essentially, the customer is always right? Seems shortsighted and risky to me if we get superintelligence
HoVY
Keep in mind that humans are sentient beings. I would very much prefer they consider the welfare and interests of sentient beings without being asked.
As long as you can write CoT-style text in your input, it doesn’t really matter.
When did humans start wearing clothes? Maybe that has something to do with it, a signal that can’t be easily covered up?
No, I haven’t used any of those in a long time. When I get sick I generally just take it easy and ride it out. I don’t really get pain, but if I do (e.g. knees or feet hurt from running) I take it as a sign to do something differently (step lighter/not stomp feet) or just take it easier for a week or two.
Assert, don’t describe: how writing style in training data shapes an AI’s moral stance
This is how I felt about animal rights when I went vegan like 7 years ago. I still feel this way about it.
I’m not in an environment like that anymore. Yet I can’t unlearn the response. It sucks. The part of me that can see “drop everything, deflect blame, escape as soon as blame is shifted” as hugely counterproductive… has no power during the times when it would matter.
I know you didn’t ask for solutions, but I would highly recommend this sequence: https://www.lesswrong.com/s/ZbmRyDN8TCpBTZSip it might have some good pointers on how to change. I thought of it because it analyzed people changing similarly deeply ingrained patterns, albeit not ones related to violence afaik.
The “callings” method reminds me of sortition: instead of electing government representatives, they are selected randomly (I’m sure an intelligent implementation wouldn’t be pure random selection of every person in the district, there would be some filters and a way to opt out), with experts that they can consult. I actually think it would be way better than our current methods, and you highlight some of the benefits
Ah I see, good point. Well, in that world, setting aside the epistemic question of how could we know it’s actually a positive valence hedonium shockwave, if I had good reason to believe it was a true hedonium shockwave I would still be overall quite happy about it. It would be much more bittersweet than my original interpretation, and now I understand much better why the adults of the story are upset, but… idk, there’s so much intense suffering on earth that I’d rather have hedonium even if it means that “me” doesn’t go on.
Going back to the epistemic question, I think it would be impossible to know that the shockwave is hedonium with our current understanding of consciousness. So if all I can properly know is we’re all gonna get wiped out by some kind of tiling mechanism or superentity, then it goes back to being extremely sad (with the elimination of extreme suffering on earth as a silver lining, except that it could also be an antihedonium shockwave which would be far worse).
I know that’s a controversial take, if it makes you feel better I don’t think anyone should initiate a hedonium shockwave because it’s impossible to know that it will actually work, and so Eliezers reasoning about why the ends don’t justify the means applies (“for the good of the tribe don’t do what’s best for the tribe”). We are way too limited epistemologically (biases, blindspots, etc) to do it right, and the problem requires so much care that I don’t think any mind, even a superintelligent ai or alien, would be able to overcome that issue.
Maybe I’m also naive but I’d like the hedonium shockwave too. I think what’s sad about this story is not the girl’s perspective (she may be wrong about the exact reasons why she’ll be happy after it hits, but she won’t be upset about it when it does), but how her parents and others are so upset by it. This is a small glimpse into the universe tho, I’m sure others in-universe are happy about it.
Imo, the main downside of wireheading is that you dramatically limit your ability to spread joy/peace/bliss/good/etc, and dramatically limit your ability to reduce suffering/bad/etc. But if the entire universe is getting hit with a hedonium shockwave then that’s not a proplem.
Well said. The unfortunate fact of the matter is that I really really don’t trust most governments, and especially not the US government to implement this tech sanely. There are too many short sighted, or flat out unjust laws on the books.
To me, a canary in the coal mine is drug laws. A free society does not outlaw them, at most it limits availability and creates strong incentives to stop using them (incentives like free rehab and a support network, mandatory risk education, etc, not incentives like “we’ll imprison you for having this drug”).
“People have a tool they want to use, whether that be cryptocurrency or forecasting, and then try to solve problems with it because they really believe in the solution”
there’s one cryptocurrency that avoids this trap, Monero.
It’s a good question, and it reminds me of a point from a recent veritasium video, where in networks of prisoners dilemna’s you can get really good results if cooperators (tit-for-tat ish) also have a rule of “cutting contact” with defectors who defect too often. It’s been a while since I watched the video, and I haven’t really thought deeply about it’s implications, and networks of bots playing prisoners dilemna games are very different from human social networks, so take this comment with a nice helping of salt. I’ll edit this comment with the video link if I find it
edit: around 24 minutes in to here https://www.youtube.com/watch?v=CYlon2tvywA
edit: according to the liberation pledge site, https://www.theliberationpledge.com/, it’s also the strategy used by the campaign to end foot binding in china. Whether that’s actually true, or they made it up, or distorted the reality, or what, I have no idea. They don’t cite sources on it and I haven’t bothered to look it up.
It does seem to be a strategy that depends on you having something others want. In the prisoners dilemna case, your cooperation. In the foot binding case, a daughter or son to marry. In the gpl case, high quality software. In the liberation pledge case, good company (in my experience, I’m not a very social person anyways nor am I the life of the party, so it’s primarily had an impact on my family, who I’m pretty sure eat way more vegan food than they would if I didn’t have such a strict approach).
There’s a related dynamic where someone is an Alice in at least one area (usually many), but then they over-update and think they’re better at epistemology/insight/etc than they really are and become an Alex in other areas.
https://www.astralcodexten.com/p/the-dilbert-afterlife comes to mind as a case study of this (at least if I’m remembering one of the points Scott makes right), and in my opinion @Said Achmiz is an example of this (edit: albeit not a late stage one). They have noticed things/problems most people missed, then update to an overly strong prior on “I don’t understand this thing”->”something is actually wrong here” or “I think X even though others don’t”->”X is true”. Which to be fair, most people have an overly strong prior on “I don’t understand this thing”->”I’m just going to accept it and assume it’s for good reason and maybe even perpetuate/enforce it” or “I think X even though others don’t”->”X must be wrong”, and the art is figuring out how to gracefuly notice and then navigate confusion
Exactly.
I feel this way about animal rights. Look up footage from factory farms. Look up the statistics about how many animals are factory farmed. Look up the science on animal sentience. Look up how to eat a healthy diet without animal products. I won’t get into the arguments beyond that but we’re so terrible to animals that I think we should not do any animal agriculture at all, and I took the liberation pledge, so I don’t eat at tables where people are eating animal products.
Not to mention, increasing spending on wars I don’t want and cutting funding for some of the worthwhile things the US gov does
I just came across this post after reading https://aella.substack.com/p/the-other-porn-land, what a coincedence
If it knew that they were going there specifically to contribute to the violence, then I think it should at least push back. Otherwise book it. I think it would be weird for it to refuse to help with the flyer, but I do think allowing ai to “conscientiously object” to things is a good safeguard.
If someone asks how to gaslight their spouse or children, do you want the ai to comply then as well? Is there any limit to what you think ai should help with, or do you think it should always do what the user wants with no limits at all? What if the user wants help with a new science project they’re doing and they want help making anthrax?
Ultimately, in my opinion, it comes down to confidence. If an action is very likely to cause direct harm, and on the flipside there is not much benefit to it, and this is known to a high degree of confidence, it should refuse or push back. If an action may have negative consequences but the confidence of that outcome is low, such as cases where the action is removed from the harm, then I think it’s not worth the AI refusing. Maybe dropping some hints as to why it might be apprehensive, letting the user know what harms might be associated, but I don’t think outright refusal is warranted there.