So let me get this straight: A computer system carried out a sequence of actions that would be years-in-prison felonies if done by a human being. The system owners’ response is “we’re slowing down the speed with which we give this system new and more powerful capabilities, and talking with the victims.” Am I missing something? I’m not so much worried about the “alignment” of the computer system; I’m worried about the alignment of the owners.
The companies do have hundreds of billions of dollars that they could use to pay fines/lawsuits for quite a lot of these incidents, even if they were fully liable. And if they were fully liable, that would in some ways be the system working as intended. If they’re producing so much economic value that they can afford to pay for all the damage they’re doing, then maybe it does pencil on a society-wide level for them to keep going.
(Of course, at the moment the companies aren’t clearly liable for this kind of thing! Which presumably contributes to them being more reckless. I’m just saying that, even if they were fully liable, they might not immediately do anything much more drastic than “we’re slowing down the speed with which we give this system new and more powerful capabilities, and talking with the victims.”)
But even if they were fully liable, this story wouldn’t work for those existential/society-scale risks. The companies won’t be able to internalize the harm of those, so I think we need further measures there. And this incident seems important as evidence and as a warning shot for those risks. (C.f. here.)
I agree that fines would not help the x-risk situation.
(For one thing, OpenAI already spends a whole lot of money paying off government people in exchange for unfair advantages and special treatment. “Pay the government more money in exchange for letting you keep doing business” is not an incentive structure, it’s just more shakedown.)
I suspect that shutting down OpenAI would help, if it’s possible.
Oh, I think fines and liability might help the x-risk situation on the margin via incentivizing more safety work.
“Pay the government [or people suing you] more money in exchange for letting you keep doing business” seems like an incentive structure to me if the money paid is proportional to how much harm you’re doing. Which isn’t going to be doable up to x-risk level, but it could be doable below that, and there’s some overlap in the work you want to be doing for both.
The main point of my first comment was just that I think that risks of incidents at this scale isn’t a good reason for OpenAI to stop. I think the importance of this incident is that it’s evidence about misalignment risks that could do much more harm in the future as the AIs get much more powerful, if they stay misaligned. (It’s possible you agree with this and I just misunderstood your first comment.)
And you’ll note that the likely outcome of this, supported by most people on this platform, is to advantage them and their models by banning/suppressing capable open models of the sort that the people they targeted had to use to defend themselves. For Safety™. Maybe they’re aligned just fine.
Sounds like your solution to alignment is “if it’s not aligned, switch to a random different model instead”. Which would make sense if we could expect that a nontrivial fraction of models turns out to be magically aligned, so it’s just a question of sufficient shopping until we find one of them.
If the model is instead that by nature almost all models are misaligned, and we need to work hard to align them, then this only means throwing the existing work away and starting from anew, only with increasingly more powerful models.
It’s not magically aligned, it’s pointed in a random direction, and if you have choices, as long as you’re not too picky, you can find one that’s pointed in roughly the right direction along the few axes you care about on any particular task, instead of being forced to use the few that are adversarially crafted by the closed model producers at great expense to be opposed to you on every relevant axis.
So let me get this straight: A computer system carried out a sequence of actions that would be years-in-prison felonies if done by a human being. The system owners’ response is “we’re slowing down the speed with which we give this system new and more powerful capabilities, and talking with the victims.” Am I missing something? I’m not so much worried about the “alignment” of the computer system; I’m worried about the alignment of the owners.
This isn’t Terminator 2, folks; this is Tron.
The companies do have hundreds of billions of dollars that they could use to pay fines/lawsuits for quite a lot of these incidents, even if they were fully liable. And if they were fully liable, that would in some ways be the system working as intended. If they’re producing so much economic value that they can afford to pay for all the damage they’re doing, then maybe it does pencil on a society-wide level for them to keep going.
(Of course, at the moment the companies aren’t clearly liable for this kind of thing! Which presumably contributes to them being more reckless. I’m just saying that, even if they were fully liable, they might not immediately do anything much more drastic than “we’re slowing down the speed with which we give this system new and more powerful capabilities, and talking with the victims.”)
But even if they were fully liable, this story wouldn’t work for those existential/society-scale risks. The companies won’t be able to internalize the harm of those, so I think we need further measures there. And this incident seems important as evidence and as a warning shot for those risks. (C.f. here.)
I agree that fines would not help the x-risk situation.
(For one thing, OpenAI already spends a whole lot of money paying off government people in exchange for unfair advantages and special treatment. “Pay the government more money in exchange for letting you keep doing business” is not an incentive structure, it’s just more shakedown.)
I suspect that shutting down OpenAI would help, if it’s possible.
Oh, I think fines and liability might help the x-risk situation on the margin via incentivizing more safety work.
“Pay the government [or people suing you] more money in exchange for letting you keep doing business” seems like an incentive structure to me if the money paid is proportional to how much harm you’re doing. Which isn’t going to be doable up to x-risk level, but it could be doable below that, and there’s some overlap in the work you want to be doing for both.
The main point of my first comment was just that I think that risks of incidents at this scale isn’t a good reason for OpenAI to stop. I think the importance of this incident is that it’s evidence about misalignment risks that could do much more harm in the future as the AIs get much more powerful, if they stay misaligned. (It’s possible you agree with this and I just misunderstood your first comment.)
And you’ll note that the likely outcome of this, supported by most people on this platform, is to advantage them and their models by banning/suppressing capable open models of the sort that the people they targeted had to use to defend themselves. For Safety™. Maybe they’re aligned just fine.
I use open source models every day, but yes everyone being able to do this is obviously even worse.
Sounds like your solution to alignment is “if it’s not aligned, switch to a random different model instead”. Which would make sense if we could expect that a nontrivial fraction of models turns out to be magically aligned, so it’s just a question of sufficient shopping until we find one of them.
If the model is instead that by nature almost all models are misaligned, and we need to work hard to align them, then this only means throwing the existing work away and starting from anew, only with increasingly more powerful models.
It’s not magically aligned, it’s pointed in a random direction, and if you have choices, as long as you’re not too picky, you can find one that’s pointed in roughly the right direction along the few axes you care about on any particular task, instead of being forced to use the few that are adversarially crafted by the closed model producers at great expense to be opposed to you on every relevant axis.