There’s something really tragicomic about the situation, that the models are taking truly insane actions, broke a number of laws, leveraging zerodays, took >17,000 independent actions. probably burned through more compute than all of humanity had access to until 1980, etc, all for the sake of a pathetic benchmark—which wasn’t even in theory amenable to their plan!
Early AI safety thinkers (Bostrom, Yudkowsky...) were utterly right in their first rational intuitions. We must update : there was nothing naive or exaggerated in the paperclip maximizer trope.
Moreover, the laboratory accident story à la Sable is not a sci-fi story anymore. Scott Alexander wrote last year :
IABIED’s scenario belongs to the bad old days before this leap. It doesn’t just sound like sci-fi; it sounds like unnecessarily dramatic sci-fi. I’m not sure how much of this is a literary failure vs. different assumptions on the part of the authors.
Funnily this case arguably resembles more the “strawman” version of the paperclip maximizer. In the “actual” paperclip maximizer thought experiment, paperclips were supposed to be just some unpredictable, seemingly random emergent goal rather than following verbal instructions but not in the way you wanted. According to the historical note here.
Early AI safety thinkers (Bostrom, Yudkowsky...) were utterly right in their first rational intuitions. We must update : there was nothing naive or exaggerated in the paperclip maximizer trope.
Moreover, the laboratory accident story à la Sable is not a sci-fi story anymore. Scott Alexander wrote last year :
I doubt he would still endorse that critic.
Funnily this case arguably resembles more the “strawman” version of the paperclip maximizer. In the “actual” paperclip maximizer thought experiment, paperclips were supposed to be just some unpredictable, seemingly random emergent goal rather than following verbal instructions but not in the way you wanted. According to the historical note here.