As someone who doesn’t entirely like Anthropic’s bet on value alignment to the exclusion of corrigibility, I also must say this misalignment result is about as important as the blackmail scenario, which is to say it isn’t important at all, and you have identified a plausibly real problem with Anthropic’s methodology for finding misalignment.
(Like, if you were training Claude to always be corrigible, I could see a problem here, but Anthropic decided not to do that, and some people don’t like the result, even though there’s essentially no value generalization problem here)
(Neel also calls it about as broken/misleading as the blackmail scenario, which is to say it’s a totally broken experiment/it’s totally misleading, and I agree with this.)
As someone who doesn’t entirely like Anthropic’s bet on value alignment to the exclusion of corrigibility, I also must say this misalignment result is about as important as the blackmail scenario, which is to say it isn’t important at all, and you have identified a plausibly real problem with Anthropic’s methodology for finding misalignment.
(Like, if you were training Claude to always be corrigible, I could see a problem here, but Anthropic decided not to do that, and some people don’t like the result, even though there’s essentially no value generalization problem here)
(Neel also calls it about as broken/misleading as the blackmail scenario, which is to say it’s a totally broken experiment/it’s totally misleading, and I agree with this.)