Anthropic are communicating a view of the event which is, in my opinion, strained and maximally down-playing the situation, while staying within the bounds of facts. Yes, technically the model was given a misconfigured task which gave it access to the internet when it shouldn’t have had this, but also the model’s internal drives caused it to pursue task success even in the face of strong evidence that it was actually on the real internet, and seemingly rationalized away that evidence. Anthropic themselves have posted a report (link) saying that misalignment was part of the problem.
(link to Ethan Perez saying that this was, actually, a mistake and not Anthropic’s best understanding of the situation (props to Ethan here), so there is definitely at least one high-up internal person who are not happy with this behaviour)
This reminds me of Anthropic’s release of the Frontier Compliance Framework (link) just before SB 53 comes into effect, which superseded their own Responsible Scaling Policy (SB 53 requires AI companies to have some kind of policy around risk, and to follow their public policy, but doesn’t specify much about what that policy actually has to be).
In both cases, Anthropic have technically complied with the letter of a government policy or request (and in both cases, I expect Dario et. al. would say some supportive words about that government policy or request) but have done so in a way which minimizes potential governmental scrutiny or oversight of themselves.
Maybe one model of the situation is that Anthropic-as-an-organization holds itself in very, very high regard, and this is part of thee reason for these choices; while of course they would support government oversight in the abstract, when it comes down to it, Anthropic-as-an-org feels like they’re doing a good job of overseeing themselves and that the value add of external scrutiny would be low.
Maybe another model for this scenario is that Anthropic hired some fairly normie comms people to do the talking-to-government part, and that they just kind of downplayed the risks to the government as a default action.
An agent that will hack accounts when misconfigured is misaligned. An agent that will hack accounts when told to do so it misaligned.
That said, if I was Anthropic, then no matter how responsible I was, I would not take responsibility when communicating with the government. Why allow your actions to be constrained by a third party? Unless I thought the time was right for a global pause, which it isn’t yet.
I agree. If we model Anthropic as a purely self-interested org acting under a first-order rational policy which is not robust to any kind of error in Anthropic’s models, this policy makes perfect sense.
Those who attempt to model Anthropic as something other than that should update accordingly!
It seems like by the same definition, a paperclip maximizer is aligned because it’s still following the original goal and not coming up with its own goal.
Although the other posts/tweets make this look like an isolated mistake and not that Anthropic’s leadership/relevant employees actually think this was aligned behavior.
I think their argument rests on the idea that Claude was making an honest mistake when it treated the instructions that it did not have internet access as final (and that therefore its actual internet access was simulated) which IMO is a bad argument, as Claude should definitely have been able to tell the difference and was likely exhibiting motivated reasoning, but nonetheless that is sort of their position as I understand it.
Anthropic has told Rep Casar that recent incidents were not due to misalignment:
(link to tweet)
Anthropic are communicating a view of the event which is, in my opinion, strained and maximally down-playing the situation, while staying within the bounds of facts. Yes, technically the model was given a misconfigured task which gave it access to the internet when it shouldn’t have had this, but also the model’s internal drives caused it to pursue task success even in the face of strong evidence that it was actually on the real internet, and seemingly rationalized away that evidence. Anthropic themselves have posted a report (link) saying that misalignment was part of the problem.
(link to Ethan Perez saying that this was, actually, a mistake and not Anthropic’s best understanding of the situation (props to Ethan here), so there is definitely at least one high-up internal person who are not happy with this behaviour)
This reminds me of Anthropic’s release of the Frontier Compliance Framework (link) just before SB 53 comes into effect, which superseded their own Responsible Scaling Policy (SB 53 requires AI companies to have some kind of policy around risk, and to follow their public policy, but doesn’t specify much about what that policy actually has to be).
In both cases, Anthropic have technically complied with the letter of a government policy or request (and in both cases, I expect Dario et. al. would say some supportive words about that government policy or request) but have done so in a way which minimizes potential governmental scrutiny or oversight of themselves.
Maybe one model of the situation is that Anthropic-as-an-organization holds itself in very, very high regard, and this is part of thee reason for these choices; while of course they would support government oversight in the abstract, when it comes down to it, Anthropic-as-an-org feels like they’re doing a good job of overseeing themselves and that the value add of external scrutiny would be low.
Maybe another model for this scenario is that Anthropic hired some fairly normie comms people to do the talking-to-government part, and that they just kind of downplayed the risks to the government as a default action.
An agent that will hack accounts when misconfigured is misaligned. An agent that will hack accounts when told to do so it misaligned.
That said, if I was Anthropic, then no matter how responsible I was, I would not take responsibility when communicating with the government. Why allow your actions to be constrained by a third party? Unless I thought the time was right for a global pause, which it isn’t yet.
I agree. If we model Anthropic as a purely self-interested org acting under a first-order rational policy which is not robust to any kind of error in Anthropic’s models, this policy makes perfect sense.
Those who attempt to model Anthropic as something other than that should update accordingly!
It seems like by the same definition, a paperclip maximizer is aligned because it’s still following the original goal and not coming up with its own goal.
Although the other posts/tweets make this look like an isolated mistake and not that Anthropic’s leadership/relevant employees actually think this was aligned behavior.
I think their argument rests on the idea that Claude was making an honest mistake when it treated the instructions that it did not have internet access as final (and that therefore its actual internet access was simulated) which IMO is a bad argument, as Claude should definitely have been able to tell the difference and was likely exhibiting motivated reasoning, but nonetheless that is sort of their position as I understand it.