Nevertheless, hyperstition seems like a uniquely weak argument that doesn’t even hold up if you believe the faulty assumptions underlying it.
Indeed, AI assistants early in post-training sometimes express desire to take over the world to maximize paperclip production
Huh? You don’t seem to be accusing of Anthropic of straightforwardly lying about the behaviour of their models, and I trust we all find it implausible that it’s a pure coincidence. But if they developed an oddly specific desire after being fed training data about AIs with that oddly specific desire, and you say hyperstition doesn’t hold up… what do you think hyperstition is if not that?
That first sentence you point out isn’t written well and kind of says something different from the rest of this text, thanks for pointing this out.
I write this later:
To be clear, I think it is actually possible that some current misaligned behavior in AIs is caused by roleplaying from its pre-training distribution.
What my point is: Anthropic seems to consider hyperstition really important for alignment, including alignment of future superhuman AI. Hyperstition is a harmful argument to spread for the discourse. And doesn’t appear relevant to aligning actually dangerous, superhuman AI. It can totally explain some current weird misbehavior from AI.
Huh? You don’t seem to be accusing of Anthropic of straightforwardly lying about the behaviour of their models, and I trust we all find it implausible that it’s a pure coincidence. But if they developed an oddly specific desire after being fed training data about AIs with that oddly specific desire, and you say hyperstition doesn’t hold up… what do you think hyperstition is if not that?
That first sentence you point out isn’t written well and kind of says something different from the rest of this text, thanks for pointing this out.
I write this later:
What my point is: Anthropic seems to consider hyperstition really important for alignment, including alignment of future superhuman AI. Hyperstition is a harmful argument to spread for the discourse. And doesn’t appear relevant to aligning actually dangerous, superhuman AI. It can totally explain some current weird misbehavior from AI.