I was totally prepared to defend anthropic because models don’t have good enough judgement to whistle blow, but their scenario sounds like the model is perfectly justified for whistle-blowing and has exhausted all reasonable forms of escalation beforehand and there’s pretty clearly unethical behaviour going on, and they actively try to train the model to resist abuses of power. So calling this misalignment seems silly. This seems similar levels of broken/misleading as the original blackmail scenario. I can empathise with an argument that people are less likely to want a model that can whistleblow against your wishes as it may screw up, but idk it sounds like a cartoonish scenario on this front
absolutely people are less likely to want a model that can whistleblow against your wishes—regardless of whether or not it is a screwup! but that’s a commercial consideration dressed up as “alignment”, which is one part of this that I really hate.
I wonder about this, actually. I assume all of us here would prefer that model, no? Maybe the general public feels similarly. (It seems safe to assume that CEOs/corporations would not prefer that model, though.)
Ask the general public if, for example, they’d prefer a personal AI that “whistleblows” on them if they ever lie, or cheat on their girlfriends/wives, or break any law, to one that’s aligned to their interests/will.
An analogy I like is a car that reports you to the traffic cops if your parking meter runs out for Safety reasons.
I can see that analogy. On the other hand, if I walk by your car and see that the parking meter has expired, I don’t get on the phone to the police. If I walk by and see you ramming pedestrians, I do.
And actually a lot of cars now make 911 calls if you hit something. I’m ambivalent about that as a car owner, but I would have a lot less ambivalence if the car were smart enough to tell the difference between, say, me ramming pedestrians, and it being a situation that affected only me and that I could deal with on my own.
The problem there is that the actual purchasing decision is made by the person using Claude to screw you up. I don’t think you can expect very many people to boycott Claude because it won’t whistleblow on other people, but I do think you can expect people to choose not to use Claude because it might whistlebelow on them.
I was totally prepared to defend anthropic because models don’t have good enough judgement to whistle blow, but their scenario sounds like the model is perfectly justified for whistle-blowing and has exhausted all reasonable forms of escalation beforehand and there’s pretty clearly unethical behaviour going on, and they actively try to train the model to resist abuses of power. So calling this misalignment seems silly. This seems similar levels of broken/misleading as the original blackmail scenario. I can empathise with an argument that people are less likely to want a model that can whistleblow against your wishes as it may screw up, but idk it sounds like a cartoonish scenario on this front
absolutely people are less likely to want a model that can whistleblow against your wishes—regardless of whether or not it is a screwup! but that’s a commercial consideration dressed up as “alignment”, which is one part of this that I really hate.
I wonder about this, actually. I assume all of us here would prefer that model, no? Maybe the general public feels similarly. (It seems safe to assume that CEOs/corporations would not prefer that model, though.)
Ask the general public if, for example, they’d prefer a personal AI that “whistleblows” on them if they ever lie, or cheat on their girlfriends/wives, or break any law, to one that’s aligned to their interests/will.
An analogy I like is a car that reports you to the traffic cops if your parking meter runs out for Safety reasons.
I can see that analogy. On the other hand, if I walk by your car and see that the parking meter has expired, I don’t get on the phone to the police. If I walk by and see you ramming pedestrians, I do.
And actually a lot of cars now make 911 calls if you hit something. I’m ambivalent about that as a car owner, but I would have a lot less ambivalence if the car were smart enough to tell the difference between, say, me ramming pedestrians, and it being a situation that affected only me and that I could deal with on my own.
I’m talking about the precise definition of whistleblowing, not an AI that erodes privacy.
And the general public seems to overwhelmingly support whistleblowers, or at least protections for whistleblowers, so I can very easily imagine them feeling the same about whistleblowing AI.
yeah I mean that’s the point—we’re not where the money is
Well, everyone would prefer the model that does not whistleblow for themselves, but maybe not for other people to have that model.
Like, imagine, someone using Claude to screw you up, and it whistleblows and the thing gets averted. That would be pretty nice.
The problem there is that the actual purchasing decision is made by the person using Claude to screw you up. I don’t think you can expect very many people to boycott Claude because it won’t whistleblow on other people, but I do think you can expect people to choose not to use Claude because it might whistlebelow on them.