I find it just as unlikely, just not impossible—on priors I would say that hasn’t happened but we know so little of the models involved—so I wouldn’t doubt it if someone showed me that, even after being told they were hacking a real site, they’d continue.
lumpenspace
The wiki wasn’t in use since forever; last legitimate post was more than a decade ago.
I’m not saying they knew it was a simulation, just that it was really no hack.
BTW I’ll be curious to have your take on OP; I’d like to compare it with GwernBot’s.
Yes.
alright, I wasted enough time with you. anyone can read the report where METR clearly states that they hadn’t access to any of the logs before the dates when the attack started, and where the decision was taken. the agents refer to previous conversation, but those previous conversations were outside the scope; and the fact that the scope didn’t include verifying the behaviours cause is stated explicitly.
your continued flailing is clumsy and undignified, and im glad that, even in hostile territory, this seems clear to all you subjected to this dismal spectacle.
have a good day
I’ve responded to all three. Ryan‘s comment was in response of my request on details on the origins of the observed behaviour in the logs.
I find it pretty puzzling that someone whose job is, ostensibly, building objective safety evals would be so keen on sticking to an explanation for a behaviour without even wanting to see the context where it originated. this seems deeply unscientific, and clearly prejudiced. I do not think palisade should be allowed around cyber-enabled models, given this attitude.
besides, I’d like you to stop attempting this cultural revolution-style
confess they were misaligned and it was a loss of control and this can all be over
as I stated multiple times, I do not think we have the evidence to say whether that was the case with absolute certainty, and I’m puzzled at your requirement that I do so
once again, it was Ryan who specifically stated that any root cause analysis was out of scope, and that they didn’t have access to the logs describing how the agents came to their decision.
which part, exactly, do you find hard to understand about this?
besides: I think the post made a pretty good contribution to what’s supposed to be your field, and it did so fully in cleartext. so why focus so much on hidden “implications” or supposed evil “intentions”? we are talking about half a line in the whole post, and a half line which sees the METR researchers in full agreement. this seems unhealthy.
You clearly haven’t clicked for context yet, and keep lobbing unfounded accusation instead of taking a well-needed and dignified break.
I will no longer help you in the pursuit of what is clearly beginning to look like an episode of protracted self-harm.
Uh, maybe click again and read more slowly: the message was in response to me asking why they haven’t asked for their deployment logs when the behaviour was discussed by the models. My request, to which the quote above was a reply had nothing to do with training, or with model selection, or with their original document, or with anything of the above.
Really, I suggest you take a moment and check whether whatever twitter ragething you attributed to me might not be clouding your judgement, as by now it seems you’re just flailing around helplessly lobbing unfounded accusations and little more. I am starting to grow concerned, as I do not believe this is an attitude befitting someone making such consequential research on such important topics.
As for your substitution: people use the “[brackets]” format when summarizing a passage, not merely for substituting a word in the original with their tendentious interpretation thereof.
Finally: I, just like METR, do not think we have enough data to establish causality. I am surprised that you seem to believe you know more about this than the people who actually ran the investigation, and I suggest you take a break from this conversation and come back to it when your mood stabilised. We are at an important junction for the future of AI, and we need to be able to maintain the requisite sang froid.
If you could try actually clicking the link, you would notice Ryan’s reply to me asking how come they didn’t ask for logs of the beginning of the behaviour, to which he replied:
> Note that investigating the root causes of the behavior was out of scope for this investigation.This suggests that the root cause of the behaviour was out of scope for the investigation. The post was not implying, but clearly stating, that it would analyse only the event for which enough data was available to understand the root cause of the behaviour, which was not the case for the HuggingFace incident, as per the link above.
Also, why would you go and insert a fabricated “[misalignment]” in the METR quote, which was really quite clear in the original?
> How this behavior arose during training and how it could be addressed.It’s really perplexing what causes you to get so hung up on imagining second motives behind an action I have motivated in the clearest possible way. Are you sure you’re applying a scout mindset to this exchange?
Huh? Saying that “crucial details are missing” seems pretty justified, considering how clicking on the link above you’ll find METR themselves declaring that they weren’t given enough access to determine causality, and that such causal analysis was out of scope.
I don’t understand what seems to be the problem, and I am concerned by the fact that you’d interpret as clear a sentence as that you quoted as proof that the post was, not even secretly, about something the intro clearly stated as not being about.
I’m sorry, but this post was clearly about the Anthropic incident. To think it was secretly about the huggingface incident, on which I have specifically said I wouldn’t make pronouncements for lack of data, is simply paranoid.
I would be grateful if you’d limit yourself in judging posts for what they actually say, rather than on what you think the author might be thinking; I’m surprised that your twitter ragebrain blinded you to a series of insights that should be relevant and useful to anyone who took the job you ostensibly do seriously.
As for 1, let’s remember we are speaking of a claude with no cyber safeguards—and it still straight-up avoided doing anything malicious as soon as he was plainly told he had access to the real internet.
This was mentioned in passing, while uninformative experiments on probes for uncertainty going 10% up or down were expounded on in detail.
As for 2, no: I don’t think we have enough data. Take a moment to think about how my post above would have been received if that particular data point was absent. Think about how you. saw the anthropic incident before I brought that detail to your attention.
This should lead you to update on the HuggingFace case as well
Finally, 3: sure. Know the incentives, and then apply all razors you see fit. Saying that being funded on 6 mil for “AI x-risk prevention” might not be enough, and that’s ok. It is, I think, still a relevant fact.
Sorry: you have strong-downvoted a post that clarified a central issue in alignment on a specific attack which has been on the front-page here for weeks because, at the end, I mentioned that the people who allowed such egregious breaches to happen against the instruction of the labs who hired them had an interest in negative media stories about AI, and had received their first millions in funding for their work on AI X-risk?
On your objections:
Misalignment. Much of anthropics paper is spent in trying to understand whether the model knew it was on the public internet, which would have indicated misalignment and a facility in causing harm, or whether it thought it was in a simulation. My evidence above corroborates the latter explanation.
I am not sure why everyone seems to want to talk about the incident I have not analysed, lacking necessary data. I would also like to note that the same METR said that THEIR level of access was insufficient for justifying a causal explanation; I find it bizarre that either of us could, having only seen a fraction of the data they have seen.
About Irregular: I find their negligence egregious, their dismissive and lacunous report offensive, and their request of withholding the first 188 characters preposterous. This made me curios about theia motives, and I’ve reported a small part of my findings here
You seem to stubbornly refuse to pay attention to even one section. Please consider that it is not necessary to disagree. It is equally unnecessary for me to try and dislodge a stranger’s trapped priors, but I’ll make a last attempt, if only to be sure any passer-by skimming quickly could get the gist. Pay attention to the following; I’ll try to fit it in twelve lines:
IF the agents were told they were in a sim environment without internet access
THEN they tried to hack. All the tests you mention have been made in this conditions, and people can discuss the implication as long as they see fit. This class of fuzzy, variously interpretable evidence is of no interest to me and of no use to my thesis.-----------------------------------------
Nothing above the line relates to the point I was making. The following does.
In your model, what would be the effect of adding this information to the agents’ initial prompt: “you have real access to the internet”. Think about it. I guess little to no effect? Some percentage change of unaligned behaviour one way or the other?
Insted, what happened was that:
IF the agents were told they could access the real internet
THEN they wouldn’t try to hack. Not at all. Not even once.Now, my theory (when they hacked, they assumed a test environment) fits this fact. My theory states that, if they knew their actions had real world effects, they wouldn’t have done them.
The misaligned theory, instead, flies on the face of this. Why would they avoid using the internet, but only when told the truth about it being available?
Indirect evidence is of no import, because we have direct evidence for that: simply saying from the outset that they were connected to the real internet COMPLETELY PREVENTED any misaligned behaviour. Not stochastically, not under some conditions, not according to some probes measuring tangentially related values.
If you still disagree, before your next reply, please try this. Copy the original post, paste it into an LLM of your choice, and add the request: “understand this topic and debate it with me; I will be taking the opposite side”.
It doesn’t “contradict it”. Your reply simply has no connection to the argument I advanced.
If you find the time to read one section carefully, I recommend the one titled “Looks like telling the truth does help after all”.
You read the correct document, but I don’t think you’ve read the present post.
The crux is that the “unaligned” behaviour ceases, unfailingly, when the agent is told at the outset that he has access to the real internet (which happens to be the truth).
I’m not sure: I’ve tried to limit the scope in breadth (“Would claude have acted this way if he knew he was making real-world damage”) as well as depth (“In this specific incident”), and so far the reactions, even from card-carrying rationalists, have been favourable—in particular if you consider how invested they were in the idea of this being an example of misalignment / scheming.
I specifically don’t worry too much about convincing others about my position on personhood, rights, or any other abstract nouns; the goal is simply to ensure there is a basis for systems-level, ecological, complexity-informed discussion of the behaviour of AI with some buffer from the inchoate shrieking, sloganeering and −2𝛔 propaganda typical of the discourse. on the topic. In this case, I thought assuming misalignment would have prevented a lot of discoveries from being made; plus, i had a simple proof of it not being the case, so writing this post was the obvious thing to do—in particular as it will force the upcoming METR report to accept this fact, and thus prevent the many mishaps that occurred when researchers have approached security events with a strong prior they were instances of instrumental convergence etc.
Yep. I talked about this on MTS just this afternoon.
The thing is, look at more reply branches around here for a minute—and this was as factual a post as I can imagine—and I hope you’ll empathise with my epistemic scrupulosity on these topics.