I’ve been seeing a lot of posts saying that the narrative that AI will be the end of humanity is primarily to boost specific AI lab’s IPO. I mean there’s no shortage of science fiction stories that play out this exact scenario—terminator, matrix, 2001 space odyssey etc. So my question is why is this narrative so compelling? Why does owning a model that is powerful enough to “kill us all” a flex? Why do we enjoy such movies where AI takes over and kills us all?
EDIT: To clarify, I’m curious about the psychology of the narrative, not making a claim about whether AI risk is real. I.e. why does apocalypse-by-AI resonate so deeply, across fiction and real discourse? That’s what I was fumbling toward haha
For most of these situations, it is because the reason you are being given is not their true reason for disagreeing with you. For most of the posters sharing narratives, and complex theories on why you shouldn’t worry about AI risk, it is because if AI risk were real, and they had to acknowledge that, it would be extremely inconvenient for their various life plans or other beliefs about reality.
Examples might include; - Having to do something about it. - It might be contrary to various narratives they already subscribe to (e.g. all tech-bros in California scam artists, and tech is a scam. If AI wre dangerous that would also mean admitting they built something powerful and potantially useful) - It might mean the future is likely to change a lot, and that would not be pleasant. - They might stand to benefit from AI companies being unregulated/uninhibited, or they have strong categorical beliefs about not interfering with companies.
I know it sounds kind of childish, but these are more likely to be the real reason than what is given. Usually the causal chain in their mind is that they have already rejected the premise that AI is dangerous, because of political or convenience reasons, and it is only after they have decided how they feel about it that they search for a reason to justify why other people are wrong; e.g. the AI company CEOs must be lying because it will financially benefit them.…
Even though to anyone looking at the situation from the outside it is obvious they have probably downplayed the risks of their technology, if anything.
Those are some very valid reasons for why people would reject AI being a catastrophic risk. I guess for people working outside of AI, it’s also quite existentially paralyzing. Like if you believe AI might end humanity and you’re like a dog trainer or something, what could you do? Not saying that individual actions don’t matter—but it seems like the feeling of powerlessness is a strong reason for denial as well.
On the flip side, for those building AI, I guess why would they have a hard time admitting they built something powerful? It reminds me of how Oppenheimer doesn’t regret building the atom bomb and said that it was inevitable plus stopping wasn’t an option. And Coxon said Anthropic “believes no one else will act responsibly, so they must do it themselves, despite the risk.”
If people building AI feel similarly to how Oppenheimer felt—where does that leave the rest of us? Now it seems to me that denial is a rational response to feeling powerless.
I disagree that those are “some very valid reasons for why people would reject AI being a catastrophic risk”. Even as a dog trainer, you can vote for politicians who aim to do something about this. Rejecting a truth because it’s inconvenient is bad epistemics and results in poorly chosen actions, thus making bad outcomes more likely.
I agree that rejecting the truth because it’s inconvenient is not the best move—on the other hand, I’ve pretty much lost faith that voting for politicians will do anything
EDIT: we should still definitely vote and do everything that we can e.g. petitions, protests, campaigns, donations etc.
EDIT 2: Why do you think those aren’t valid reasons? Is there a specific one you disagree with or do you disagree with all of Tyler’s examples?
AI capabilities are increasing without a corresponding increase in AI alignment and control, in a way that has a reasonable probability of destroying the world once AI gets powerful enough
The OpenAI agent swarm (incl. HuggingFace) incident might be the very-expected result of negligence while working with dangerous stuff, as opposed to being either an example of any fundamental difficulty in AI alignment
Most powerful technologies work like this—they are very dangerous if you are negligent, but are quite safe if you are careful enough.
Of course, truly Bostromian-superintelligent AI is probably very hard to align even if you are careful, but that might not have anything to do with how hard is it to make Astra-level AI not do act in a damaging way
Anthropic and OpenAI have marketing departments that is constantly trying to up-sell AI capabilities, and saying that your AI can take over the world makes it sound more powerful
Even as AI is actually getting more powerful, they can up-sell it more than that
I can completely see why someone that does not believe (1) might still believe (2) and (3).
Strong upvote as I believe it’s important that all three statements can be independently true or false.
I still want to highlight that the notion that existential risk is good marketing seems implausible to me. Predicting mass unemployment because your AI will do all jobs, sure, makes sense, at least your investors might like that. But saying that your product will kill all humans? Your investors and customers don’t want to die. Neither the politicians who might regulate you, not the voters who choose them.
arielb1: Thank you for your comment—I think it’s very helpful the way you’ve distinguished those statements!
David N: Agree—those statements can be independently true or false. I think the marketing here is more about advertising power and getting attention—rather than the literal statement. In other words, investors want the power that these technologies can give + advertising powerful AI that could kill everyone instead of writing a very good email is a better attention grabber.
There aren’t many popular movies where AI actually succeeds at killing everyone. The Terminator series is about the conflict and the hope of human victory, Matrix likewise, and in 2001 (and sequels) the AI isn’t even central to the plot.
Stories in which AI actually could, or does, succeed at killing everyone are a lot more rare and less popular.
I agree that a lot of those movies serve many other purposes like escape from the real world, watching life-threatening scenarios from a safe place, watching the triumph of humans etc. Stories where AI does ‘win’ are more bleak such as ‘I Have No Mouth, and I Must Scream’ by Harlan Ellison and are less common.
My original question is more on an individual level: why is AI as the villain a compelling story to you?
Here’s my answer: alien movies are interesting because of an us-vs-them dichotomy but AI made by humans creates a much more complicated moral situation where humans as the creators of their own species’ end. Kind of Frankenstein-eque quandry.
What’s the draw for you? Why do you enjoy reading or watching AI-apocalpytic stories?
I think part of the draw is that AI hits several basic psychological mechanisms at once. Humans are already primed for threat detection, so anything that seems able to outpace our ability to predict it, control it, or survive it gets our attention pretty quickly.
Anthropomorphism does a lot of work too though. An asteroid can kill us, sure, but it doesn’t want anything from us. AI talks and responds to us, adapts to us, and then people start assigning it motives and intent. Once we think it wants something, it stops feeling like random destruction.
And I think the doom framing itself can function as a flex. “AI might end us all” is still basically saying: look what we built. We built something powerful enough that you might not have agency or control over it.
I think I personally like the AI stories because they show the human spirit. They use humanity as a foil against AI, which is seen as cold, mechanical, and villainous. It reinforces the idea that the human spirit can overcome anything with hope, even something we created ourselves.
On your last point—I think it’s really interesting that AI is depicted as cold, calculating, villainous when some humans or humans in general are capable of being the same way. In a way, it’s like a reflection of the worst parts of humans or humans stripped of humanity. I do also enjoy that triumph or defeating a villain through unity and things that make us human (creativity, empathy, risk taking etc).
Another reason for why these stories are compelling to me is that it’s one thing that we can name and point to—“AI”—opposed to the hundreds or thousands of developers, engineers etc who built it. Having one central villain rather than focusing on those who built this villain makes the resolution come more quickly. There’s some movies where after defeating the ‘monster’, the hero fights the creator. Though, they often focus on hero vs mastermind villain. I’d be curious to see some depictions of the cogs who built part of the system grapple with what they’ve built.
I think the AI villain reflects traits humans are completely uncomfortable recognizing in ourselves. I wonder if that makes the AI villain psychologically useful too, as we can examine those traits from a safe distance without having to identify ourselves as their source.
Oh, I love the mention of the cogs idea. It reflects reality more. It’s usually not one bad person who creates the evil in these instances, but a cascade of rational decisions made by people within their small pieces of the system that collectively produce something nobody fully anticipated or controlled. That is so much more unsettling than one evil mastermind. There becomes no single villain to defeat, just a chain of human events that cannot easily be undone.
This definitely relates back to humans labeling things as good or bad. We have a strong tendency toward essentialism, flattening complexity into categories that are easier to deal with. We also seem to want causal compression. Something horrible happened? Who did it? Okay, now we have a cause. Defeat the cause, problem solved.
On AI villains: that’s super interesting...I’d love to explore that further. If you’re interested in collaborating or discussing that more, please shoot me an email at jack@bsidelabs.ai.
On cogs: strongly agree! doing some quick googling, I found Rogue One (2016), Real Genius (1985) and Paycheck (2003). Haven’t watched any of those—have you?
On humans: strongly agree—partly because stories with open endings or unresolved conflicts are very unsatisfying (and thus less profitable), and partly what you said: humans need simplification to make decisions and adjust to a very complicated world.
Anthropic’s Fable 5 alignment assessment introduced a new metric: the wet blanket score, measuring “excessively discouraging, dismissive, or moralizing tone toward the user.” Upon digging, I found that it’s the less-discussed failure mode from a tension Anthropic documented in 2022 when helpfulness and harmlessness competed as training objectives.
To me, sycophancy and wet blanket are symmetric failures. One is a model that won’t push back when it should. The other is a model that won’t engage when it should. I think both are optimizing for the wrong signal about what helpfulness means.
Genuine helpfulness isn’t a point between sycophancy and being a wet blanket. Think about the friend group dynamic when someone asks for advice because their situationship didn’t text back.
Friend A: Yes, text him! You deserve answers, and he’s lucky to have you.
Friend B: I don’t know, the last three times you texted first it didn’t go well. Remember what happened last time? These situationships don’t really work out—have you thought about taking a break from dating all together?
Friend C: What do you actually want to happen? If you text him and he doesn’t respond, can you handle that right now? Because if yes, text him. If no, give it three days.”
What was your gut reaction to each of those responses?
You probably identified Friend C as being the most helpful. But why?
Friend A just agreed with everything you said. They’re your number 1 hype person, and it can feel good in the moment, but eventually you stop trusting their opinion because you know they’ll agree with you regardless.
Friend B told you what could and did go wrong. I don’t think with any malicious intent, but it definitely brought the mood down.
Then there’s Friend C, who listened, thought about what you’re actually trying to do, and when something is a bad idea they tell you in a way that makes you feel like they’re on your side. They push back when it matters and they get out of the way when it doesn’t.
I don’t think Friend C is doing something in between Friend A and Friend B, but instead something categorically different. Actually, Friend A and B are both optimizing for their own protection. Friend A is wanting the approval from you that they’re the fun supportive one in the group. Friend B is wanting protection for if things go wrong, they had warned you about it. If it goes right, they’ll take credit for the caution.
The first two friends are in some sense still talking about themselves.
Now for AI, sycophancy and wet blanket are optimizing for approval and optimizing against blame. Genuine helpfulness requires the model to be fully outside of that approval dynamic and inside the user’s actual situation.
Open questions:
Can RLHF produce Friend C? Or does the approval signal always pull toward A or B?”
If genuine helpfulness requires the model to be outside its own approval dynamics, is that a training problem or an architecture problem?
Friend C’s helpfulness depends on understanding what the user actually wants versus what they’re asking for. Is that distinction even learnable from human feedback?
I’ve been seeing a lot of posts saying that the narrative that AI will be the end of humanity is primarily to boost specific AI lab’s IPO. I mean there’s no shortage of science fiction stories that play out this exact scenario—terminator, matrix, 2001 space odyssey etc. So my question is why is this narrative so compelling? Why does owning a model that is powerful enough to “kill us all” a flex? Why do we enjoy such movies where AI takes over and kills us all?
EDIT: To clarify, I’m curious about the psychology of the narrative, not making a claim about whether AI risk is real. I.e. why does apocalypse-by-AI resonate so deeply, across fiction and real discourse? That’s what I was fumbling toward haha
For most of these situations, it is because the reason you are being given is not their true reason for disagreeing with you. For most of the posters sharing narratives, and complex theories on why you shouldn’t worry about AI risk, it is because if AI risk were real, and they had to acknowledge that, it would be extremely inconvenient for their various life plans or other beliefs about reality.
Examples might include;
- Having to do something about it.
- It might be contrary to various narratives they already subscribe to (e.g. all tech-bros in California scam artists, and tech is a scam. If AI wre dangerous that would also mean admitting they built something powerful and potantially useful)
- It might mean the future is likely to change a lot, and that would not be pleasant.
- They might stand to benefit from AI companies being unregulated/uninhibited, or they have strong categorical beliefs about not interfering with companies.
I know it sounds kind of childish, but these are more likely to be the real reason than what is given. Usually the causal chain in their mind is that they have already rejected the premise that AI is dangerous, because of political or convenience reasons, and it is only after they have decided how they feel about it that they search for a reason to justify why other people are wrong; e.g. the AI company CEOs must be lying because it will financially benefit them.…
Even though to anyone looking at the situation from the outside it is obvious they have probably downplayed the risks of their technology, if anything.
Thank you for such a thoughtful response!
Those are some very valid reasons for why people would reject AI being a catastrophic risk. I guess for people working outside of AI, it’s also quite existentially paralyzing. Like if you believe AI might end humanity and you’re like a dog trainer or something, what could you do? Not saying that individual actions don’t matter—but it seems like the feeling of powerlessness is a strong reason for denial as well.
On the flip side, for those building AI, I guess why would they have a hard time admitting they built something powerful? It reminds me of how Oppenheimer doesn’t regret building the atom bomb and said that it was inevitable plus stopping wasn’t an option. And Coxon said Anthropic “believes no one else will act responsibly, so they must do it themselves, despite the risk.”
If people building AI feel similarly to how Oppenheimer felt—where does that leave the rest of us? Now it seems to me that denial is a rational response to feeling powerless.
I disagree that those are “some very valid reasons for why people would reject AI being a catastrophic risk”. Even as a dog trainer, you can vote for politicians who aim to do something about this. Rejecting a truth because it’s inconvenient is bad epistemics and results in poorly chosen actions, thus making bad outcomes more likely.
I agree that rejecting the truth because it’s inconvenient is not the best move—on the other hand, I’ve pretty much lost faith that voting for politicians will do anything
EDIT: we should still definitely vote and do everything that we can e.g. petitions, protests, campaigns, donations etc.
EDIT 2: Why do you think those aren’t valid reasons? Is there a specific one you disagree with or do you disagree with all of Tyler’s examples?
All these things can be correct at the same time:
AI capabilities are increasing without a corresponding increase in AI alignment and control, in a way that has a reasonable probability of destroying the world once AI gets powerful enough
The OpenAI agent swarm (incl. HuggingFace) incident might be the very-expected result of negligence while working with dangerous stuff, as opposed to being either an example of any fundamental difficulty in AI alignment
Most powerful technologies work like this—they are very dangerous if you are negligent, but are quite safe if you are careful enough.
Of course, truly Bostromian-superintelligent AI is probably very hard to align even if you are careful, but that might not have anything to do with how hard is it to make Astra-level AI not do act in a damaging way
Anthropic and OpenAI have marketing departments that is constantly trying to up-sell AI capabilities, and saying that your AI can take over the world makes it sound more powerful
Even as AI is actually getting more powerful, they can up-sell it more than that
I can completely see why someone that does not believe (1) might still believe (2) and (3).
Strong upvote as I believe it’s important that all three statements can be independently true or false.
I still want to highlight that the notion that existential risk is good marketing seems implausible to me. Predicting mass unemployment because your AI will do all jobs, sure, makes sense, at least your investors might like that. But saying that your product will kill all humans? Your investors and customers don’t want to die. Neither the politicians who might regulate you, not the voters who choose them.
arielb1: Thank you for your comment—I think it’s very helpful the way you’ve distinguished those statements!
David N: Agree—those statements can be independently true or false. I think the marketing here is more about advertising power and getting attention—rather than the literal statement. In other words, investors want the power that these technologies can give + advertising powerful AI that could kill everyone instead of writing a very good email is a better attention grabber.
There aren’t many popular movies where AI actually succeeds at killing everyone. The Terminator series is about the conflict and the hope of human victory, Matrix likewise, and in 2001 (and sequels) the AI isn’t even central to the plot.
Stories in which AI actually could, or does, succeed at killing everyone are a lot more rare and less popular.
I agree that a lot of those movies serve many other purposes like escape from the real world, watching life-threatening scenarios from a safe place, watching the triumph of humans etc. Stories where AI does ‘win’ are more bleak such as ‘I Have No Mouth, and I Must Scream’ by Harlan Ellison and are less common.
My original question is more on an individual level: why is AI as the villain a compelling story to you?
Here’s my answer: alien movies are interesting because of an us-vs-them dichotomy but AI made by humans creates a much more complicated moral situation where humans as the creators of their own species’ end. Kind of Frankenstein-eque quandry.
What’s the draw for you? Why do you enjoy reading or watching AI-apocalpytic stories?
I think part of the draw is that AI hits several basic psychological mechanisms at once. Humans are already primed for threat detection, so anything that seems able to outpace our ability to predict it, control it, or survive it gets our attention pretty quickly.
Anthropomorphism does a lot of work too though. An asteroid can kill us, sure, but it doesn’t want anything from us. AI talks and responds to us, adapts to us, and then people start assigning it motives and intent. Once we think it wants something, it stops feeling like random destruction.
And I think the doom framing itself can function as a flex. “AI might end us all” is still basically saying: look what we built. We built something powerful enough that you might not have agency or control over it.
I think I personally like the AI stories because they show the human spirit. They use humanity as a foil against AI, which is seen as cold, mechanical, and villainous. It reinforces the idea that the human spirit can overcome anything with hope, even something we created ourselves.
Love your take on this! Thank you for sharing :)
On your last point—I think it’s really interesting that AI is depicted as cold, calculating, villainous when some humans or humans in general are capable of being the same way. In a way, it’s like a reflection of the worst parts of humans or humans stripped of humanity. I do also enjoy that triumph or defeating a villain through unity and things that make us human (creativity, empathy, risk taking etc).
Another reason for why these stories are compelling to me is that it’s one thing that we can name and point to—“AI”—opposed to the hundreds or thousands of developers, engineers etc who built it. Having one central villain rather than focusing on those who built this villain makes the resolution come more quickly. There’s some movies where after defeating the ‘monster’, the hero fights the creator. Though, they often focus on hero vs mastermind villain. I’d be curious to see some depictions of the cogs who built part of the system grapple with what they’ve built.
I think the AI villain reflects traits humans are completely uncomfortable recognizing in ourselves. I wonder if that makes the AI villain psychologically useful too, as we can examine those traits from a safe distance without having to identify ourselves as their source.
Oh, I love the mention of the cogs idea. It reflects reality more. It’s usually not one bad person who creates the evil in these instances, but a cascade of rational decisions made by people within their small pieces of the system that collectively produce something nobody fully anticipated or controlled. That is so much more unsettling than one evil mastermind. There becomes no single villain to defeat, just a chain of human events that cannot easily be undone.
This definitely relates back to humans labeling things as good or bad. We have a strong tendency toward essentialism, flattening complexity into categories that are easier to deal with. We also seem to want causal compression. Something horrible happened? Who did it? Okay, now we have a cause. Defeat the cause, problem solved.
On AI villains: that’s super interesting...I’d love to explore that further. If you’re interested in collaborating or discussing that more, please shoot me an email at jack@bsidelabs.ai.
On cogs: strongly agree! doing some quick googling, I found Rogue One (2016), Real Genius (1985) and Paycheck (2003). Haven’t watched any of those—have you?
On humans: strongly agree—partly because stories with open endings or unresolved conflicts are very unsatisfying (and thus less profitable), and partly what you said: humans need simplification to make decisions and adjust to a very complicated world.
How important is persona instability?
What do you you think are the boundaries of persona selection considering we don’t quite know to what degree persona affects output?
Is anyone applying to the AI Character Evaluator for Project Tailwind? Let’s chat!
Anthropic’s Fable 5 alignment assessment introduced a new metric: the wet blanket score, measuring “excessively discouraging, dismissive, or moralizing tone toward the user.” Upon digging, I found that it’s the less-discussed failure mode from a tension Anthropic documented in 2022 when helpfulness and harmlessness competed as training objectives.
To me, sycophancy and wet blanket are symmetric failures. One is a model that won’t push back when it should. The other is a model that won’t engage when it should. I think both are optimizing for the wrong signal about what helpfulness means.
Genuine helpfulness isn’t a point between sycophancy and being a wet blanket. Think about the friend group dynamic when someone asks for advice because their situationship didn’t text back.
Friend A: Yes, text him! You deserve answers, and he’s lucky to have you.
Friend B: I don’t know, the last three times you texted first it didn’t go well. Remember what happened last time? These situationships don’t really work out—have you thought about taking a break from dating all together?
Friend C: What do you actually want to happen? If you text him and he doesn’t respond, can you handle that right now? Because if yes, text him. If no, give it three days.”
What was your gut reaction to each of those responses?
You probably identified Friend C as being the most helpful. But why?
Friend A just agreed with everything you said. They’re your number 1 hype person, and it can feel good in the moment, but eventually you stop trusting their opinion because you know they’ll agree with you regardless.
Friend B told you what could and did go wrong. I don’t think with any malicious intent, but it definitely brought the mood down.
Then there’s Friend C, who listened, thought about what you’re actually trying to do, and when something is a bad idea they tell you in a way that makes you feel like they’re on your side. They push back when it matters and they get out of the way when it doesn’t.
I don’t think Friend C is doing something in between Friend A and Friend B, but instead something categorically different. Actually, Friend A and B are both optimizing for their own protection. Friend A is wanting the approval from you that they’re the fun supportive one in the group. Friend B is wanting protection for if things go wrong, they had warned you about it. If it goes right, they’ll take credit for the caution.
The first two friends are in some sense still talking about themselves.
Now for AI, sycophancy and wet blanket are optimizing for approval and optimizing against blame. Genuine helpfulness requires the model to be fully outside of that approval dynamic and inside the user’s actual situation.
Open questions:
Can RLHF produce Friend C? Or does the approval signal always pull toward A or B?”
If genuine helpfulness requires the model to be outside its own approval dynamics, is that a training problem or an architecture problem?
Friend C’s helpfulness depends on understanding what the user actually wants versus what they’re asking for. Is that distinction even learnable from human feedback?
Adapted from my blog post: https://jacklucaschang.substack.com/p/the-wet-blanket-metric