It doesn’t matter how many fake versions of you hold the wrong conclusion about their own ontological status, since those fake beliefs exist in fake versions of you. The moral harm caused by a single real Chantiel thinking they’re not real is infinitely greater than infinitely many non-real Chantiels thinking they are real.
Interesting. When you say “fake” versions of myself, do you mean simulations? If so, I’m having a hard time seeing how that could be true. Specifically, what’s wrong about me thinking I might not be “real”? I mean, if I though I was in a simulation, I think I’d do pretty much the same things I would do if I thought I wasn’t in a simulation. So I’m not sure what the moral harm is.
Do you have any links to previous discussions about this?
To be clear, it does seem to me that the issues discussed in this article take up relatively little of the probability mass of bad outcomes. My main AI safety concern is still just regular value-misalignment. But I wanted to at least mention the concern in the article, since it seems like something worth at least keeping in the back of your mind.
I don’t know of any human analogue that’s as clear and extreme as literally hacking yourself. However, I think there are milder examples. For example, I used to have a rather severe case of generalized anxiety disorder. Many people with generalized anxiety disorder might realize that obsessively ruminating about unlikely ways in which things could go terribly wrong is not a helpful or productive thing to do. However, acknowledging that this is not worth doing, in many cases, does not result in people stopping the rumination. Similar analogues can be found with depressed people ruminating on everything wrong with their lives and the world, people with obsessive-compulsive disorder obsessing over their compulsions, and people with attention deficit disorder failing to focus on things even though they know they really should.
I know above I talked about people with mental illnesses, but my impression is that mentally healthy people can also suffer from the above issues sometimes, albeit in milder forms.
Naturally, you don’t see humans completely hacking their own minds, at the very least because people simply don’t know how to.
I don’t know, maybe. It doesn’t seem to me that we currently have AIs advanced enough for the concerns discussed in the article to be serious issues. And I don’t really have a good sense of what future, more capable AIs with cognitively be like. But I’m curious about the reasoning you used to arrive at this conclusion, if you’re interested in sharing.
By this, do you mean there are multiple different mappings from what we’d intuitively call a “thought” to a concrete encoding in neural weights? I didn’t manage to find much information for or against this online, but I could have very easily missed something about this.
For what it’s worth, one modicum of evidence against this is that it provides poor compression: if many different neural activations encode what’s effectively the same thought, you potentially could make the reasoning system more space-efficient by removing this redundancy. And I suppose polysemanticity suggests there’s decent amounts of optimization pressure toward space efficiency in artificial neural networks.
Edit: Also, I want to mention that in the post I talked about dangerous “thoughts”, but the argument generalizes to dangerous computational states that don’t map cleanly onto individual thoughts. For example, maybe a thought T on its own is not in general dangerous, but a specific encoding E of that thought T would be. Then the basic argument for the potential danger would be basically the same as described in the post except that the rogue AI would be trying to get Coral to think up that specific encoding E, rather just any encoding of the thought.