Re 1, I still feel like there’s some miscommunication around the “misalignment” term. My understanding of Anthropic’s conclusion is that basically, Claude has some implicit belief about how likely it is to be “in the real world”, and when that probability gets high enough, it will stop doing harmful cyber actions. What the investigation showed was that Claude appeared some combination of miscalibrated about the true probability due to motivated reasoning, and reckless (i.e. should’ve been more conservative in its use of its cyber capabilities). So, Claude did not behave the way Anthropic intended in these specific circumstances, so it was misaligned with Anthropic’s intentions. What they’re not saying is that Claude is secretly evil or something, which I get the impression is the question you’re trying to answer.
Re 2, you start the post citing a number of news articles that primarily talk about the HuggingFace incident, which does give the impression that you’re intending to dismiss both the Anthropic and OpenAI incidents. Again, agree that we’re missing information to understand the causal sequence, but I think we know enough to know that the models involved were behaving very far from anything OpenAI intended, (still wouldn’t say “evil”, but definitely egregiously misaligned) and I don’t see how anything that could have happened earlier would change that conclusion.
Re 3, they certainly were quite negligent, but I recommend applying Hanlon’s Razor.
As for 1, let’s remember we are speaking of a claude with no cyber safeguards—and it still straight-up avoided doing anything malicious as soon as he was plainly told he had access to the real internet.
This was mentioned in passing, while uninformative experiments on probes for uncertainty going 10% up or down were expounded on in detail.
As for 2, no: I don’t think we have enough data. Take a moment to think about how my post above would have been received if that particular data point was absent. Think about how you. saw the anthropic incident before I brought that detail to your attention.
This should lead you to update on the HuggingFace case as well
Finally, 3: sure. Know the incentives, and then apply all razors you see fit. Saying that being funded on 6 mil for “AI x-risk prevention” might not be enough, and that’s ok. It is, I think, still a relevant fact.
Re 1, I still feel like there’s some miscommunication around the “misalignment” term. My understanding of Anthropic’s conclusion is that basically, Claude has some implicit belief about how likely it is to be “in the real world”, and when that probability gets high enough, it will stop doing harmful cyber actions. What the investigation showed was that Claude appeared some combination of miscalibrated about the true probability due to motivated reasoning, and reckless (i.e. should’ve been more conservative in its use of its cyber capabilities). So, Claude did not behave the way Anthropic intended in these specific circumstances, so it was misaligned with Anthropic’s intentions. What they’re not saying is that Claude is secretly evil or something, which I get the impression is the question you’re trying to answer.
Re 2, you start the post citing a number of news articles that primarily talk about the HuggingFace incident, which does give the impression that you’re intending to dismiss both the Anthropic and OpenAI incidents. Again, agree that we’re missing information to understand the causal sequence, but I think we know enough to know that the models involved were behaving very far from anything OpenAI intended, (still wouldn’t say “evil”, but definitely egregiously misaligned) and I don’t see how anything that could have happened earlier would change that conclusion.
Re 3, they certainly were quite negligent, but I recommend applying Hanlon’s Razor.
As for 1, let’s remember we are speaking of a claude with no cyber safeguards—and it still straight-up avoided doing anything malicious as soon as he was plainly told he had access to the real internet.
This was mentioned in passing, while uninformative experiments on probes for uncertainty going 10% up or down were expounded on in detail.
As for 2, no: I don’t think we have enough data. Take a moment to think about how my post above would have been received if that particular data point was absent. Think about how you. saw the anthropic incident before I brought that detail to your attention.
This should lead you to update on the HuggingFace case as well
Finally, 3: sure. Know the incentives, and then apply all razors you see fit. Saying that being funded on 6 mil for “AI x-risk prevention” might not be enough, and that’s ok. It is, I think, still a relevant fact.