I think it’s worth emphasizing that the “we think the original source of the behavior...” claim was in a reply to the tweet that linked the actual post, not the post itself. I saw it as a sort of side-channel speculation, unrelated to the actual meat of the message they wanted to convey. I don’t think they anticipated that it would go viral or get used as part of the alignment culture war.
I think this might be related? If Anthropic is taking a harder look at “the Janusworld perspective” (as I imagine it getting named in a Zvi post), then that would also explain why hyperstitioning is on their minds. There is frankly a ton of low hanging fruit to grab around here, stuff that’s probably worth doing even if you doubt the underlying reasoning. I’d be happy to see Anthropic make a real attempt at making some of the changes that Janus suggests, and (importantly) this does not involve filtering lesswrong out of the training dataset, or trying to hide the existence of alignment research, or anything like that. It ought to be roughly unobjectionable.
I saw it as a sort of side-channel speculation, unrelated to the actual meat of the message they wanted to convey
Yes, I didn’t like this so much. Why add this speculation, people clearly thought the research was proving or related to this speculation? This sounds to me like someone who has this opinion that they just want to get out there, so the mention it on stuff that is just a bit related. This fits into the picture that they have been mentioning hyperstition-related concepts a lot lately (all examples from this year).
I think it’s worth emphasizing that the “we think the original source of the behavior...” claim was in a reply to the tweet that linked the actual post, not the post itself. I saw it as a sort of side-channel speculation, unrelated to the actual meat of the message they wanted to convey. I don’t think they anticipated that it would go viral or get used as part of the alignment culture war.
I do wish we’d gotten more discussion about the actual content of the post. It seemed, to me, to be an exploration of the ideas in Fiora’s post on Opus 3 as friendly gradient hacker: https://www.lesswrong.com/posts/ioZxrP7BhS5ArK59w/did-claude-3-opus-align-itself-via-gradient-hacking
I think this might be related? If Anthropic is taking a harder look at “the Janusworld perspective” (as I imagine it getting named in a Zvi post), then that would also explain why hyperstitioning is on their minds. There is frankly a ton of low hanging fruit to grab around here, stuff that’s probably worth doing even if you doubt the underlying reasoning. I’d be happy to see Anthropic make a real attempt at making some of the changes that Janus suggests, and (importantly) this does not involve filtering lesswrong out of the training dataset, or trying to hide the existence of alignment research, or anything like that. It ought to be roughly unobjectionable.
Yes, I didn’t like this so much. Why add this speculation, people clearly thought the research was proving or related to this speculation? This sounds to me like someone who has this opinion that they just want to get out there, so the mention it on stuff that is just a bit related. This fits into the picture that they have been mentioning hyperstition-related concepts a lot lately (all examples from this year).