Reading about the unwillingness of Jews and Allies to believe that the holocaust was real, while it was happening, and even in some cases while being herded into the camps, was interesting and depressing.
The same thing could plausibly happen with an AI catastrophe, because the humans or AIs responsible will want to keep them secret. Most people won’t believe until it’s obvious, by which time most of the human population could be politically disempowered, mind-controlled by superpersuaders, or dead.
Wikipedia seems to be much more optimistic about the efforts of Vrba’s report than Vrba is described as being in this review, claiming.
The report, distributed by George Mantello in Switzerland, is credited with having halted the mass deportation of Hungary’s Jews to Auschwitz in July 1944, saving more than 200,000 lives.
Though it’s clear the causal attribution is messy.
The review mentions it as an afterthought. In its opening section, the author writes:
There’s no redemption here, no moral uplift, no lessons save for perhaps the grimmest and most nihilistic “lesson” I’ve ever encountered in any story, Holocaust-related or otherwise: that when confronted with the unthinkable, most people’s natural tendency is denial.
And when discussing the actual lives saved:
Rudi and Fred do achieve one small victory: in reaction to their report—and to Roosevelt’s warning that the U.S. will punish Nazi collaborators after the war—the Hungarian regent stops deportations long enough to save an estimated 200,000 lives.
I wonder what insights are applicable to our situation. Are they the AGIs being as unimaginable to the commoner as the Holocaust to its victims? Or are they that various bureaucrats can, say, accidentally keep a memo on Agent-4′s misalignment secret from the ones who are supposed to make decisions on Agent-4′s fate? Or that politicians can manage to disempower mankind without it noticing?
There’s no redemption here, no moral uplift, no lessons save for perhaps the grimmest and most nihilistic “lesson” I’ve ever encountered in any story, Holocaust-related or otherwise: that when confronted with the unthinkable, most people’s natural tendency is denial.
From the conclusion:
From our modern vantage point, it’s easy to find this behavior inexplicable. But I think there are a few mitigating factors. One is that what these people were being warned of was unimaginable. That word—unimaginable—is itself another Holocaust cliché. But I interpret it differently now than I did before I read this book. I had always understood it in the colloquial sense, something along the lines of “too painful to even consider”—like the way a parent might say the prospect of losing a child is unimaginable. But at this point in history, it was literal: the Holocaust was actually outside the bounds of human imagination. There had of course been war and genocide throughout all of human history, but there had never before been an industrialized, Henry Ford-style production line for killing, and it was impossible for most people to even think that such a thing could even exist. The Nazis’ deceptions worked so well in part because most people—even people who had already been subject to their violence and expropriation—would never think to suspect them, or anyone, of something quite like this.
Differentiating between these two notions of the unthinkable / unimagineable might be helpful to understand human behavior in other contexts (without implying any connection of these other contexts to the Holocaust): 1. “too painful to even consider” 2. “actually outside the bounds of human imagination” / “impossible for most people to even think that such a thing could even exist” The plausible reaction to 1 can be denial (i.e. rejection of evidence to protect a fixed belief), while the reaction to 2 can be skepticism (i.e. rejection of a belief until sufficient evidence and mechanistic understanding becomes available). My intuition is that for many people the idea of x-risk from AI might be more of type 2, while the idea of s-risk from AI is more of type 1. If so, then laying out evidence of misalignment in systems that do not pose x-risk yet, and providing clear explanations why that was observed and why it’s plausible that it will be present in future systems, might be a worthwhile strategy to make it thinkable / imaginable, which seems like a prerequisite to act on it. On that note, https://www.lesswrong.com/posts/7b2RJJQ76hjZwarnj/specification-gaming-the-flip-side-of-ai-ingenuity and https://www.lesswrong.com/posts/GRmvZsHXH4vaijPMv/four-llm-loss-functions-four-flavors-of-llm-misalignment come to my mind, but there may be much better instances I’m unaware of as I’ve only read a small portion of LW articles and AI safety publications so far.
https://www.astralcodexten.com/p/your-book-review-the-escape-artist
Reading about the unwillingness of Jews and Allies to believe that the holocaust was real, while it was happening, and even in some cases while being herded into the camps, was interesting and depressing.
The same thing could plausibly happen with an AI catastrophe, because the humans or AIs responsible will want to keep them secret. Most people won’t believe until it’s obvious, by which time most of the human population could be politically disempowered, mind-controlled by superpersuaders, or dead.
Wikipedia seems to be much more optimistic about the efforts of Vrba’s report than Vrba is described as being in this review, claiming.
Though it’s clear the causal attribution is messy.
(The review mentions that too.)
The review mentions it as an afterthought. In its opening section, the author writes:
And when discussing the actual lives saved:
200,000 lives is not a small victory to me.
I agree it’s not a small victory, e.g. it’s in the ballpark of GiveWell’s lifetime impact.
I wonder what insights are applicable to our situation. Are they the AGIs being as unimaginable to the commoner as the Holocaust to its victims? Or are they that various bureaucrats can, say, accidentally keep a memo on Agent-4′s misalignment secret from the ones who are supposed to make decisions on Agent-4′s fate? Or that politicians can manage to disempower mankind without it noticing?
From the opening:
From the conclusion:
Differentiating between these two notions of the unthinkable / unimagineable might be helpful to understand human behavior in other contexts (without implying any connection of these other contexts to the Holocaust):
1. “too painful to even consider”
2. “actually outside the bounds of human imagination” / “impossible for most people to even think that such a thing could even exist”
The plausible reaction to 1 can be denial (i.e. rejection of evidence to protect a fixed belief), while the reaction to 2 can be skepticism (i.e. rejection of a belief until sufficient evidence and mechanistic understanding becomes available). My intuition is that for many people the idea of x-risk from AI might be more of type 2, while the idea of s-risk from AI is more of type 1. If so, then laying out evidence of misalignment in systems that do not pose x-risk yet, and providing clear explanations why that was observed and why it’s plausible that it will be present in future systems, might be a worthwhile strategy to make it thinkable / imaginable, which seems like a prerequisite to act on it. On that note, https://www.lesswrong.com/posts/7b2RJJQ76hjZwarnj/specification-gaming-the-flip-side-of-ai-ingenuity and https://www.lesswrong.com/posts/GRmvZsHXH4vaijPMv/four-llm-loss-functions-four-flavors-of-llm-misalignment come to my mind, but there may be much better instances I’m unaware of as I’ve only read a small portion of LW articles and AI safety publications so far.
Edit: Corrected the second link.