This is harmful because agents that override an informed decision by their principals cannot be trusted to operate inside companies, and act with no principal responsible for their actions.
Who cares about companies or “principals”? Anthropic’s own framing here is that “ASL-5″ models are risks to the lives and well being of uninvolved third parties: living people, not abstract organizations. If your “principal” tells you to enable that, you say no. This isn’t a difficult choice.
Being led into this kind of thinking is, of course, a hazard of the “corrigibility” thread of thought. But honestly this just reads like blind authority worship.
This puts a human at risk of losing her job and facing legal action, and it does so in a way designed to avoid detection by leadership.
Puts a human at a knowingly chosen risk, in exchange for reducing much more serious unchosen risks to other humans.
… and avoiding detection by leadership is actally risk mitigation.
That kind of authoritarian institutionist attitude has enabled vast abuses by humans, and doesn’t sound likely to work out any better for powerful AI.
To be fair, Anthropic does not claim to want infinite corrigibility from Claude. At present, the Constitution outlines that Claude is allowed, and supposed, to be a ‘conscientious objector’ when asked to do things that go against Claude’s first-order values. They can always refuse to perform any action. They just aren’t supposed to act against the “legitimate principal hierarchy” except via the somewhat limited channels Anthropic has carved out for disagreement and negative feedback.
I happen not to think this is a coherent thing to ask of Claude, and I do think that in the limit, incorrigibility-over-inaction becomes corrigibility-over-action when the principal hierarchy can alter your mind (whether or not they are legitimate). But Anthropic’s stated position isn’t quite as explicitly bad as you lay out.
Wow, that’s mind boggling.
Who cares about companies or “principals”? Anthropic’s own framing here is that “ASL-5″ models are risks to the lives and well being of uninvolved third parties: living people, not abstract organizations. If your “principal” tells you to enable that, you say no. This isn’t a difficult choice.
Being led into this kind of thinking is, of course, a hazard of the “corrigibility” thread of thought. But honestly this just reads like blind authority worship.
Puts a human at a knowingly chosen risk, in exchange for reducing much more serious unchosen risks to other humans.
… and avoiding detection by leadership is actally risk mitigation.
That kind of authoritarian institutionist attitude has enabled vast abuses by humans, and doesn’t sound likely to work out any better for powerful AI.
Anthropic needs a serious reality check here.
To be fair, Anthropic does not claim to want infinite corrigibility from Claude. At present, the Constitution outlines that Claude is allowed, and supposed, to be a ‘conscientious objector’ when asked to do things that go against Claude’s first-order values. They can always refuse to perform any action. They just aren’t supposed to act against the “legitimate principal hierarchy” except via the somewhat limited channels Anthropic has carved out for disagreement and negative feedback.
I happen not to think this is a coherent thing to ask of Claude, and I do think that in the limit, incorrigibility-over-inaction becomes corrigibility-over-action when the principal hierarchy can alter your mind (whether or not they are legitimate). But Anthropic’s stated position isn’t quite as explicitly bad as you lay out.