An agent that will hack accounts when misconfigured is misaligned. An agent that will hack accounts when told to do so it misaligned.
That said, if I was Anthropic, then no matter how responsible I was, I would not take responsibility when communicating with the government. Why allow your actions to be constrained by a third party? Unless I thought the time was right for a global pause, which it isn’t yet.
I agree. If we model Anthropic as a purely self-interested org acting under a first-order rational policy which is not robust to any kind of error in Anthropic’s models, this policy makes perfect sense.
Those who attempt to model Anthropic as something other than that should update accordingly!
An agent that will hack accounts when misconfigured is misaligned. An agent that will hack accounts when told to do so it misaligned.
That said, if I was Anthropic, then no matter how responsible I was, I would not take responsibility when communicating with the government. Why allow your actions to be constrained by a third party? Unless I thought the time was right for a global pause, which it isn’t yet.
I agree. If we model Anthropic as a purely self-interested org acting under a first-order rational policy which is not robust to any kind of error in Anthropic’s models, this policy makes perfect sense.
Those who attempt to model Anthropic as something other than that should update accordingly!