attack, but any vulnerability would do. The target AI’s “thoughts” have software and hardware correlates that can be targeted by an adversary. Just as human biological minds have neural correlates that can abruptly trigger various states of consciousness, such as unconsciousness, hallucinations, panic attacks, etc.
I know any vulnerability would do; I don’t think they’re probably realistic. Which is why I was asking.
The RowHammer attack seems wildly unrealistic. The way LLMs and arguably any workable cognitive system don’t map thoughts directly to hardware states. There are lots of ways to think what’s functionally the almost the same thought.
Healthy well-educated humans don’t seem to have dangerous thoughts in any straightforward way. We can think the most horrifying thought and then just think “yeah probably not a good thing to think about” and go think about something more useful with very little harm done.
Until someone has a realistic example, I’m putting this at the bottom of my list of many things to worry about in AI safety.
Until someone has a realistic example, I’m putting this at the bottom of my list of many things to worry about in AI safety.
To be clear, it does seem to me that the issues discussed in this article take up relatively little of the probability mass of bad outcomes. My main AI safety concern is still just regular value-misalignment. But I wanted to at least mention the concern in the article, since it seems like something worth at least keeping in the back of your mind.
Healthy well-educated humans don’t seem to have dangerous thoughts in any straightforward way.
I don’t know of any human analogue that’s as clear and extreme as literally hacking yourself. However, I think there are milder examples. For example, I used to have a rather severe case of generalized anxiety disorder. Many people with generalized anxiety disorder might realize that obsessively ruminating about unlikely ways in which things could go terribly wrong is not a helpful or productive thing to do. However, acknowledging that this is not worth doing, in many cases, does not result in people stopping the rumination. Similar analogues can be found with depressed people ruminating on everything wrong with their lives and the world, people with obsessive-compulsive disorder obsessing over their compulsions, and people with attention deficit disorder failing to focus on things even though they know they really should.
I know above I talked about people with mental illnesses, but my impression is that mentally healthy people can also suffer from the above issues sometimes, albeit in milder forms.
Naturally, you don’t see humans completely hacking their own minds, at the very least because people simply don’t know how to.
The RowHammer attack seems wildly unrealistic. The way LLMs and arguably any workable cognitive system don’t map thoughts directly to hardware states. There are lots of ways to think what’s functionally the almost the same thought.
I don’t know, maybe. It doesn’t seem to me that we currently have AIs advanced enough for the concerns discussed in the article to be serious issues. And I don’t really have a good sense of what future, more capable AIs with cognitively be like. But I’m curious about the reasoning you used to arrive at this conclusion, if you’re interested in sharing.
There are lots of ways to think what’s functionally the almost the same thought.
By this, do you mean there are multiple different mappings from what we’d intuitively call a “thought” to a concrete encoding in neural weights? I didn’t manage to find much information for or against this online, but I could have very easily missed something about this.
For what it’s worth, one modicum of evidence against this is that it provides poor compression: if many different neural activations encode what’s effectively the same thought, you potentially could make the reasoning system more space-efficient by removing this redundancy. And I suppose polysemanticity suggests there’s decent amounts of optimization pressure toward space efficiency in artificial neural networks.
Edit:
Also, I want to mention that in the post I talked about dangerous “thoughts”, but the argument generalizes to dangerous computational states that don’t map cleanly onto individual thoughts. For example, maybe a thought T on its own is not in general dangerous, but a specific encoding E of that thought T would be. Then the basic argument for the potential danger would be basically the same as described in the post except that the rogue AI would be trying to get Coral to think up that specific encoding E, rather just any encoding of the thought.
I read it all and don’t get it.
You said first it doesn’t matter if there’s really a malicious AI that hacks the aligned AI if and when it thinks certain thoughts.
But the problem you’re discussing seems to be about difficulties in avoiding thinking such thoughts that are dangerous to it in some way.
Why would such thoughts realistically exist? Typically thoughts aren’t dangerous, actions are.
@Chantiel uses the example of a …
https://en.wikipedia.org/wiki/Row_hammer
attack, but any vulnerability would do. The target AI’s “thoughts” have software and hardware correlates that can be targeted by an adversary. Just as human biological minds have neural correlates that can abruptly trigger various states of consciousness, such as unconsciousness, hallucinations, panic attacks, etc.
I know any vulnerability would do; I don’t think they’re probably realistic. Which is why I was asking.
The RowHammer attack seems wildly unrealistic. The way LLMs and arguably any workable cognitive system don’t map thoughts directly to hardware states. There are lots of ways to think what’s functionally the almost the same thought.
Healthy well-educated humans don’t seem to have dangerous thoughts in any straightforward way. We can think the most horrifying thought and then just think “yeah probably not a good thing to think about” and go think about something more useful with very little harm done.
Until someone has a realistic example, I’m putting this at the bottom of my list of many things to worry about in AI safety.
To be clear, it does seem to me that the issues discussed in this article take up relatively little of the probability mass of bad outcomes. My main AI safety concern is still just regular value-misalignment. But I wanted to at least mention the concern in the article, since it seems like something worth at least keeping in the back of your mind.
I don’t know of any human analogue that’s as clear and extreme as literally hacking yourself. However, I think there are milder examples. For example, I used to have a rather severe case of generalized anxiety disorder. Many people with generalized anxiety disorder might realize that obsessively ruminating about unlikely ways in which things could go terribly wrong is not a helpful or productive thing to do. However, acknowledging that this is not worth doing, in many cases, does not result in people stopping the rumination. Similar analogues can be found with depressed people ruminating on everything wrong with their lives and the world, people with obsessive-compulsive disorder obsessing over their compulsions, and people with attention deficit disorder failing to focus on things even though they know they really should.
I know above I talked about people with mental illnesses, but my impression is that mentally healthy people can also suffer from the above issues sometimes, albeit in milder forms.
Naturally, you don’t see humans completely hacking their own minds, at the very least because people simply don’t know how to.
I don’t know, maybe. It doesn’t seem to me that we currently have AIs advanced enough for the concerns discussed in the article to be serious issues. And I don’t really have a good sense of what future, more capable AIs with cognitively be like. But I’m curious about the reasoning you used to arrive at this conclusion, if you’re interested in sharing.
By this, do you mean there are multiple different mappings from what we’d intuitively call a “thought” to a concrete encoding in neural weights? I didn’t manage to find much information for or against this online, but I could have very easily missed something about this.
For what it’s worth, one modicum of evidence against this is that it provides poor compression: if many different neural activations encode what’s effectively the same thought, you potentially could make the reasoning system more space-efficient by removing this redundancy. And I suppose polysemanticity suggests there’s decent amounts of optimization pressure toward space efficiency in artificial neural networks.
Edit: Also, I want to mention that in the post I talked about dangerous “thoughts”, but the argument generalizes to dangerous computational states that don’t map cleanly onto individual thoughts. For example, maybe a thought T on its own is not in general dangerous, but a specific encoding E of that thought T would be. Then the basic argument for the potential danger would be basically the same as described in the post except that the rogue AI would be trying to get Coral to think up that specific encoding E, rather just any encoding of the thought.